Abstract P336: Assistance From Automated ASPECTS Software Improves Reader Performance
Bibliographic record
Abstract
Purpose: To compare physicians’ ability to read Alberta Stroke Program Early CT Score (ASPECTS) in patients with a large vessel occlusion within 6 hours of symptom onset when assisted by a machine learning-based automatic software tool, RAPID ASPECTS, compared with their unassisted score. Materials and Methods: 50 baseline CT scans selected from two prior studies (CRISP and GAMES-RP) were read by 3 experienced neuroradiologists who were provided access to a follow-up MRI. The average ASPECT score of these reads was used as the reference standard. Two additional neuroradiologists and 6 non-neuroradiologist readers then read the scans both with and without assistance from the RAPID ASPECTS software and reader improvement was determined. The primary hypothesis was that the agreement between typical readers and the consensus of 3 expert neuroradiologists would be improved with RAPID-assisted vs. unassisted reads. Agreement was based on the percentage of the individual ASPECT regions (50 cases, 10 regions each; N=500) where agreement was achieved. Results: Typical non-neuroradiologist readers agreed with the expert consensus read in 72% of the 500 ASPECTS regions, evaluated without software assistance. The automated software alone agreed in 77%. When the typical readers read the scan in conjunction with the software, agreement improved to 78% (P<0.0001, test of proportions). RAPID ASPECTS alone achieved correlations for total ASPECT scores that were similar to the expert readers who had access to the follow-up MRI scan to help enhance the quality of their reads. Conclusion: Typical readers had statistically significant improvement in their scoring of scans when the scan was read in conjunction with the automated RAPID ASPECTS software, achieving agreement rates that were comparable to neuroradiologists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".