Multicenter Automated Central Vein Sign Detection Performs as Well as Manual Assessment for the Diagnosis of Multiple Sclerosis
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: The central vein sign (CVS) is a proposed diagnostic imaging biomarker for multiple sclerosis (MS). The proportion of white matter lesions exhibiting the CVS (CVS+) is higher in patients with MS compared with its radiologic mimics. Evaluation for CVS+ lesions in prior studies has been performed by manual rating, an approach that is time-consuming and has variable interrater reliability. Accurate automated methods would facilitate efficient assessment for CVS. The objective of this study was to compare the performance of an automated CVS detection method with manual rating for the diagnosis of MS. MATERIALS AND METHODS: 3T MRI was acquired in 86 participants undergoing evaluation for MS in a 9-site multicenter study. Participants presented with either typical or atypical clinical syndromes for MS. An automated CVS detection method was employed and compared with manual rating, including total CVS+ proportion and a simplified counting method in which experts visually identified up to 6 CVS+ lesions by using FLAIR* contrast (a voxelwise product of T2 FLAIR and postcontrast T2*-EPI). RESULTS: Automated CVS processing was completed in 79 of 86 participants (91%), of whom 28 (35%) fulfilled the 2017 McDonald criteria at the time of imaging. The area under the receiver operating characteristic curve (AUC) for discrimination between participants with and without MS for the automated CVS approach was 0.78 (95% CI: [0.67,0.88]). This was not significantly different from simplified manual counting methods (select6*) (0.80 [0.69,0.91]) or manual assessment of total CVS+ proportion (0.89 [0.82,0.96]). In a sensitivity analysis excluding 11 participants whose MRI exhibited motion artifact, the AUC for the automated method was 0.81 [0.70,0.91], which was not statistically different from that for select6* (0.79 [0.68,0.92]) or manual assessment of total CVS+ proportion (0.89 [0.81,0.97]). CONCLUSIONS: Automated CVS assessment was comparable to manual CVS scoring for differentiating patients with MS from those with other diagnoses. Large, prospective, multicenter studies utilizing automated methods and enrolling the breadth of disorders referred for suspicion of MS are needed to determine optimal approaches for clinical implementation of an automated CVS detection method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".