Evaluation of Intra‐ and Interscanner Reliability of MRI Protocols for Spinal Cord Gray Matter and Total Cross‐Sectional Area Measurements
Bibliographic record
Abstract
Background In vivo quantification of spinal cord atrophy in neurological diseases using MRI has attracted increasing attention. Purpose To compare across different platforms the most promising imaging techniques to assess human spinal cord atrophy. Study Type Test/retest multiscanner study. Subjects Twelve healthy volunteers. Field Strength/Sequence Three different 3T scanner platforms (Siemens, Philips, and GE) / optimized phase sensitive inversion recovery (PSIR), T 1 ‐weighted (T 1 ‐w), and T 2 *‐weighted (T 2 *‐w) protocols. Assessment On all images acquired, two operators assessed contrast‐to‐noise ratio (CNR) between gray matter (GM) and white matter (WM), and between WM and cerebrospinal fluid (CSF); one experienced operator measured total cross‐sectional area (TCA) and GM area using JIM and the Spinal Cord Toolbox (SCT). Statistical Tests Coefficient of variation (COV); intraclass correlation coefficient (ICC); mixed effect models; analysis of variance ( t ‐tests). Results For all the scanners, GM/WM CNR was higher for PSIR than T 2 *‐w ( P < 0.0001) and WM/CSF CNR for T 1 ‐w was the highest ( P < 0.0001). For TCA, using JIM, median COVs were smaller than 1.5% and ICC >0.95, while using SCT, median COVs were in the range 2.2–2.75% and ICC 0.79–0.95. For GM, despite some failures of the automatic segmentation, median COVs using SCT on T 2 *‐w were smaller than using JIM manual PSIR segmentations. In the mixed effect models, the subject was always the main contributor to the variance of area measurements and scanner often contributed to TCA variance ( P < 0.05). Using JIM, TCA measurements on T 2 *‐w were different than on PSIR ( P = 0.0021) and T 1 ‐w ( P = 0.0018), while using SCT, no notable differences were found between T 1 ‐w and T 2 *‐w ( P = 0.18). JIM and SCT‐derived TCA were not different on T 1 ‐w ( P = 0.66), while they were different for T 2 *‐w ( P < 0.0001). GM area derived using SCT/T 2 *‐w versus JIM/PSIR were different ( P < 0.0001). Data Conclusion The present work sets reference values for the magnitude of the contribution of different effects to cord area measurement intra‐ and interscanner variability. Level of Evidence : 1 Technical Efficacy : Stage 4 J. Magn. Reson. Imaging 2019;49:1078–1090.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".