ASPECTS Interobserver Agreement of 100 Investigators from the TENSION Study
Bibliographic record
Abstract
PURPOSE: Evaluating the extent of cerebral ischemic infarction is essential for treatment decisions and assessment of possible complications in patients with acute ischemic stroke. Patients are often triaged according to image-based early signs of infarction, defined by Alberta Stroke Program Early CT Score (ASPECTS). Our aim was to evaluate interrater reliability in a large group of readers. METHODS: We retrospectively analyzed 100 investigators who independently evaluated 20 non-contrast computed tomography (NCCT) scans as part of their qualification program for the TENSION study. Test cases were chosen by four neuroradiologists who had previously scored NCCT scans with ASPECTS between 0 and 8 and high interrater agreement. Percent and interrater agreements were calculated for total ASPECTS, as well as for each ASPECTS region. RESULTS: Percent agreements for ASPECTS ratings was 28%, with interrater agreement of 0.13 (95% confidence interval, CI 0.09-0.16), at zero tolerance allowance and 66%, with interrater agreement of 0.32 (95% CI: 0.21-0.44), at tolerance allowance set by TENSION inclusion criteria. ASPECTS region with highest level of agreement was the insular cortex (percent agreement = 96%, interrater agreement = 0.96 (95% CI: 0.94-0.97)) and with lowest level of agreement the M3 region (percent agreement = 68%, interrater agreement = 0.39 [95% CI: 0.17-0.61]). CONCLUSION: Interrater agreement reliability for total ASPECTS and study enrollment was relatively low but seems sufficient for practical application. Individual region analysis suggests that some are particularly difficult to evaluate, with varying levels of reliability. Potential impairment of the supraganglionic region must be examined carefully, particularly with respect to the decision whether or not to perform mechanical thrombectomy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".