Volumetric Analysis from a Harmonized Multisite Brain MRI Study of a Single Subject with Multiple Sclerosis
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: MR imaging can be used to measure structural changes in the brains of individuals with multiple sclerosis and is essential for diagnosis, longitudinal monitoring, and therapy evaluation. The North American Imaging in Multiple Sclerosis Cooperative steering committee developed a uniform high-resolution 3T MR imaging protocol relevant to the quantification of cerebral lesions and atrophy and implemented it at 7 sites across the United States. To assess intersite variability in scan data, we imaged a volunteer with relapsing-remitting MS with a scan-rescan at each site. MATERIALS AND METHODS: All imaging was acquired on Siemens scanners (4 Skyra, 2 Tim Trio, and 1 Verio). Expert segmentations were manually obtained for T1-hypointense and T2 (FLAIR) hyperintense lesions. Several automated lesion-detection and whole-brain, cortical, and deep gray matter volumetric pipelines were applied. Statistical analyses were conducted to assess variability across sites, as well as systematic biases in the volumetric measurements that were site-related. RESULTS: < .01 for both T1 and T2 lesion volumes), with site explaining >90% of the variation (range, 13.0-16.4 mL in T1 and 15.9-20.1 mL in T2) in lesion volumes. Site also explained >80% of the variation in most automated volumetric measurements. Output measures clustered according to scanner models, with similar results from the Skyra versus the other 2 units. CONCLUSIONS: Even in multicenter studies with consistent scanner field strength and manufacturer after protocol harmonization, systematic differences can lead to severe biases in volumetric analyses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".