Multiple sclerosis cortical lesion detection with deep learning at ultra-high-field MRI
Bibliographic record
Abstract
Abstract Manually segmenting multiple sclerosis (MS) cortical lesions (CL) is extremely time-consuming, and past studies have shown only moderate inter-rater reliability. To accelerate this task, we developed a deep learning-based framework (CLAIMS: Cortical Lesion Artificial Intelligence-based assessment in Multiple Sclerosis) for the automated detection and classification of MS CL with 7T MRI. Two 7T datasets, acquired at different sites, were considered. The first consisted of 60 scans that include 0.5mm isotropic MP2RAGE acquired 4 times (MP2RAGEx4), 0.7mm MP2RAGE, 0.5mm T2*-weighted GRE, and 0.5mm T2*-weighted EPI. The second dataset consisted of 20 scans including only 0.75×0.75×0.9 mm MP2RAGE. CLAIMS was first evaluated using 6-fold cross-validation with single and multi-contrast 0.5mm MRI input. Second, performance of the model was tested on 0.7mm MP2RAGE images after training with either 0.5mm MP2RAGEx4, 0.7mm MP2RAGE, or alternating the two. Third, its generalizability was evaluated on the second external dataset and compared with a state-of-the-art technique based on partial volume estimation and topological constraints (MSLAST). CLAIMS trained only with MP2RAGEx4 achieved comparable results to the multi-contrast model, reaching a CL true positive rate of 74% with a false positive rate of 30%. Detection rate was excellent for leukocortical and subpial lesions (83%, and 70%, respectively), whereas it reached 53% for intracortical lesions. The correlation between disability measures and CL count was similar for manual and CLAIMS lesion counts. Applying a domain-scanner adaptation approach and testing CLAIMS on the second dataset, the performance was superior to MSLAST when considering a minimum lesion volume of 6μL (lesion-wise detection rate of 71% vs 48%). The proposed framework outperforms previous state-of-the-art methods for automated CL detection across scanners and protocols. In the future, CLAIMS may be useful to support clinical decisions at 7T MRI, especially in the field of diagnosis and differential diagnosis of multiple sclerosis patients.
Stored with the screening record, where it is evidence for the labels above.
How this classification was reachedexpand
The three-model screen
all 5,600 screened works →All three models called this out of scope.
Deep learning framework for detecting MS cortical lesions; a clinical imaging method, and the reliability language is domain measurement, not the reproducibility literature.
The work develops a deep-learning tool for detecting multiple-sclerosis lesions.
Deep-learning tool for detecting MS cortical lesions is clinical imaging methods development, not study of research methods.
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".