MétaCan
Menu
Back to cohort
Record W4405971752 · doi:10.46475/asean-jr.v25i3.901

Reliability and radiologists’ concordance of artificial intelligence (AI)-calculated Alberta Stroke Program Early CT Score (ASPECTS)

2025· article· en· W4405971752 on OpenAlexaboutno aff
Warissara Kiththiworaphongkich, Nuttamon Khamwongsa, Pranruethai Chaimongkol

Bibliographic record

VenueThe ASEAN Journal of Radiology · 2025
Typearticle
Languageen
FieldMedicine
TopicAcute Ischemic Stroke Management
Canadian institutionsnot available
Fundersnot available
KeywordsConcordanceReliability (semiconductor)Medical physicsComputer scienceArtificial intelligenceMedicinePsychologyReliability engineeringEngineeringInternal medicine

Abstract

fetched live from OpenAlex

Abstract Background: ASPECTS was developed for the semi-quantitative assessment of early ischemic changes (EIC) on non-contrast computed tomography (NCCT) in acute ischemic stroke (AIS). Artificial intelligence (AI)-based automated tools for the ASPECT scoring system were developed to automate the diagnosis and improve the agreement with radiologists of AIS. The performance of the automated software compared to physicians should be tested before the software is further used in clinical practice as a tool for clinicians. Objective: To evaluate the agreement with radiologists of an AI-based automated post-processing software for detecting EIC and calculating ASPECTS on NCCT images in AIS patients using a radiologist's assessment as a reference. Materials and Methods: NCCT of AIS patients were retrospectively reviewed (Stroke Fast Track Service July 2022 - December 2023). The complete set of clinical data and imaging data from both baseline and follow-up were analyzed by a radiologist as a reference. Two additional observers provided individual ASPECTS from the baseline NCCT only (observer 1 was a radiologist who independently reviewed only the baseline NCCT with stroke window setting. Observer 2 was a radiologist on service which was from the pool of 20 radiologists onsite and online). Recon&GO Inline ASPECTS software (Somaris X, VA40A, Siemens Healthineers AG, Erlangen, Germany) was applied. Both ASPECT score analysis and ASPECTS region analysis were evaluated. Positive percent agreement (PPA) and negative percent agreement (NPA) were calculated. Interobserver agreement was assessed using the Cohen's kappa coefficient and the intraclass correlation coefficient (ICC). Results: 111 patients with a mean age of 67.8 years (±11.9), 56 (50.5%) females, a mean National Institute of Health Stroke Scale (NIHSS) score of 14.2 (±8.8), and a mean time to baseline NCCT of 123.9 minutes (±58.7) were included. For dichotomized ASPECTS, the automated software showed lower PPA (14.6% vs. 27.1%) but higher NPA (100.0% vs. 93.7%) than observer 2. For the region-based analysis, both the automated software and observer 2 differed in terms of regional contribution. The automated software showed low PPA but rather high NPA with perfect (100%) NPA in lentiform nucleus and M2. The automated software showed higher agreement with the reference and two observers in deep/central regions than cortical regions. For total ASPECTS, the automated software showed a moderate agreement of total ASPECTS with the reference and observer 1 (ICC = 0.545 and 0.545). Observer 2 showed a poor agreement of total ASPECTS with the reference, observer 1, and the automated software (ICC = 0.349, 0.422, and 0.301, respectively). Conclusion: For total ASPECT score, the agreement of the tested AI software is lower compared to observer 1 obtained by a radiologist using the stroke window on NCCT, but better compared to a pool of radiologists on service with a time limit of 30 minutes to interpret the ASPECT score. When analyzing the ASPECTS regions, there are different advantages for the assessment of the deep regions and the cortical regions. The tested AI software shows higher agreement in deep/central regions than cortical regions. From the result, the tested AI software retains its potential for use in emergency situations, particularly for radiologists with limited experience and limited time to report.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.341
Threshold uncertainty score0.469

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.303
Teacher spread0.284 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueThe ASEAN Journal of RadiologySame topicAcute Ischemic Stroke ManagementFrench-language works237,207