MRI Diffusion‐Weighted Imaging to Measure Infarct Volume: Assessment of Manual Segmentation Variability
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: Manual segmentation of infarct volume on follow-up MRI diffusion-weighted imaging (MRI-DWI) is considered the gold standard but is prone to rater variability. We assess the variability of manual segmentations of MRI-DWI infarct volume. METHODS: Consecutive patients (May 2018 to May 2019) with the anterior circulation stroke and endovascularly treated were enrolled. All patients underwent 24- to 32-hour follow-up MRI. Three users manually segmented DWI infarct volumes slice by slice twice. The reference standard of DWI infarct volume was generated by the STAPLE algorithm. Intra- and interrater reliability was evaluated using the intraclass correlation coefficient (ICC) by comparing manual segmentations with the reference standard. Spatial measurements were evaluated using metrics of the Dice similarity coefficient (DSC). Volumetric measurements were compared using the lesion volume. RESULTS: The dataset consisted of 44 patients, mean (SD) age was 70.1 years (±10.3), 43% were women, and median baseline NIHSS score was 16. Among three users, the mean DSC for MRI-DWI infarct volume segmentations ranged from 80.6% ± 11.7% to 88.6% ± 7.5%, and the mean absolute volume difference was 2.8 ± 6.8 to 13.0 ± 14.0 ml. Interrater ICC among the users for DSC and infarct volume was .86 (95% confidence interval [95% CI]: .78-.91) and .997 (95% CI: .995-.998). Intrarater ICC for the three users was .83 (95% CI: .69-.93), .84 (95% CI: .72-.91), and .80 (95% CI: .64-.89) for DSC, and .99 (95% CI: .987-.996), .991 (95% CI: .983-.995), and .996 (95% CI: .993-.998) for infarct volume. CONCLUSIONS: Manual segmentation of infarct volume on follow-up MRI-DWI shows excellent agreement and good spatial overlap with the reference standard, suggesting its usefulness for measuring infarct volume on 24- to 32-hour MRI-DWI.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".