Interobserver agreement between senior radiology resident, neuroradiology fellow, and experienced neuroradiologist in the rating of Alberta Stroke Program Early Computed Tomography Score (ASPECTS)
Bibliographic record
Abstract
PURPOSE: The distribution of ischemic changes caused by infarction of the middle cerebral artery (MCA) territories is usually measured using the Alberta Stroke Program Early Computed Tomography Score (ASPECTS). The first interpreter of the brain computed tomography (CT) in the emergency department is the on-call radiology resident. The primary objective of this study was to describe the agreement of the ASPECTS performed retrospectively by the resident compared with expert raters. The second objective was to ascertain the appropriate window setting for early detection of acute ischemic stroke and good interobserver agreement between the interpreters. METHODS: We identified consecutive patients presenting with hemiparesis or aphasia at the emergency department who underwent brain CT and CT angiography. Each scan was rated using ASPECTS by senior radiology resident, neuroradiology fellow, and later by consensus between two expert raters. Statistical analysis included determination of Cohen's kappa (κ) coefficient and intraclass correlation coefficient (ICC). RESULTS: A total of 43 patients met our study criteria. Interobserver agreements for ASPECTS varied from 0.486 to 0.678 in Cohen's κ coefficient between consensus of two neuroradiologists and a neuroradiology fellow, and from 0.198 to 0.491 for consensus between two neuroradiologists and a senior radiology resident. ICC among three raters (expert consensus, neuroradiology fellow, and senior radiology resident), was very good when 8 HU window width and 32 HU center level setting was used. CONCLUSION: ASPECTS varied among raters. However, when using a narrowed window setting for interpretation, interobserver agreement improved.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.029 | 0.064 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".