FLAIR2 improves LesionTOADS automatic segmentation of multiple sclerosis lesions in non-homogenized, multi-center, 2D clinical magnetic resonance images
Bibliographic record
Abstract
Accurate segmentation of MS lesions on MRI is difficult and, if performed manually, time consuming. Automatic segmentations rely strongly on the image contrast and signal-to-noise ratio. Literature examining segmentation tool performances in real-world multi-site data acquisition settings is scarce. FLAIR2, a combination of T2-weighted and fluid attenuated inversion recovery (FLAIR) images, improves tissue contrast while suppressing CSF. We compared the use of FLAIR and FLAIR2 in LesionTOADS, OASIS and the lesion segmentation toolbox (LST) when applied to non-homogenized, multi-center 2D-imaging data. Lesions were segmented on 47 MS patient data sets obtained from 34 sites using LesionTOADS, OASIS and LST, and compared to a semi-automatically generated reference. The performance of FLAIR and FLAIR2 was assessed using the relative lesion volume difference (LVD), Dice coefficient (DSC), sensitivity (SEN) and symmetric surface distance (SSD). Performance improvements related to lesion volumes (LVs) were evaluated for all tools. For comparison, LesionTOADS was also used to segment lesions from 3 T single-center MR data of 40 clinically isolated syndrome (CIS) patients. Compared to FLAIR, the use of FLAIR2 in LesionTOADS led to improvements of 31.6% (LVD), 14.0% (DSC), 25.1% (SEN), and 47.0% (SSD) in the multi-center study. DSC and SSD significantly improved for larger LVs, while LVD and SEN were enhanced independent of LV. OASIS showed little difference between FLAIR and FLAIR2, likely due to its inherent use of T2w and FLAIR. LST replicated the benefits of FLAIR2 only in part, indicating that further optimization, particularly at low LVs is needed. In the CIS study, LesionTOADS did not benefit from the use of FLAIR2 as the segmentation performance for both FLAIR and FLAIR2 was heterogeneous. In this real-world, multi-center experiment, FLAIR2 outperformed FLAIR in its ability to segment MS lesions with LesionTOADS. The computation of FLAIR2 enhanced lesion detection, at minimally increased computational time or cost, even retrospectively. Further work is needed to determine how LesionTOADS and other tools, such as LST, can optimally benefit from the improved FLAIR2 contrast.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".