Abstract WP177: Impact of Imaging Acquisition Protocol on Automated ASPECTS Performance
Bibliographic record
Abstract
Introduction: Automated imaging analysis tools are increasingly used in clinical decision-making for stroke. Rapid ASPECTS (iSchemaView, Menlo Park, CA) assists physicians by automatically calculating Alberta Stroke Program Early CT Scores (ASPECTS) and reducing inter-reader variability. To understand why the tool’s performance in real-world settings sometimes varies compared to published literature, we investigated how different imaging acquisition protocols affect its performance. Materials&Methods: Consecutive code stroke NCCT scans with thin (1.25 mm slice; 0.625 mm spacing) and thick (5.0 mm slice; 3 mm spacing) series were collected from a retrospective database between February 2020 and May 2021. Ground truth ASPECTS reads were collected from radiology reports, which neuroradiologists determined in real-time. Automated reads were obtained using Rapid ASPECTS 1.0 and 3.0 (iSchemaView, Menlo Park, CA). Agreement between automated and manual reads was defined as ASPECTS scores within two points. Results: A total of 682 cases were included in this analysis. 67 cases were excluded for technical inadequacy (hemorrhages, tumors, and artifacts). A review of the source imaging revealed that many cases had thick overlapping slices and incorrect head positioning (neck extended instead of the standard neutral position). These cases required significant tilt correction to align the patient data with the Rapid ASPECTS regions template. These corrections led to partial voluming artifacts, which caused lower Hounsfield unit (HU) values and ASPECTS scores. When adjusting protocols from thick to thin slices, agreement between ASPECTS V1 and manual reads improved from 85% (581/682) to 89% (606/682). ASPECTS V3 showed further improvement, with agreements of 91% (619/682) and 95% (648/682) for thick and thin slice scans, respectively. Conclusion: The combination of neck extension head positioning and thick overlapping slices caused partial voluming artifacts, resulting in artificially low ASPECTS scores on automated software. Our findings indicate that adjusting imaging protocols and working with the AI provider can enhance an algorithm's accuracy. To ensure that commercially available automated analysis tools deliver accurate results, it is crucial to follow the recommended imaging acquisition protocols.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.033 | 0.134 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".