Easy-to-use and easy-to-interpret quality control of 3D gradient echo T1-weighted MR acquisition sequences for improved test-retest stability of MRI-based hippocampus volumetry
Bibliographic record
Abstract
BackgroundMRI-based hippocampus volume (HV) is widely used as neurodegeneration marker in Alzheimer's disease.ObjectiveAn easy-to-use and easy-to-interpret method to categorize T1-weighted MR sequences with respect to test-retest stability of hippocampus volumetry based on general image quality metrics (IQM).MethodsThe study included 446 3D T1-weighted MRI scans of one healthy middle-aged man obtained during 32 months in 122 scanning sessions performed with 96 different scanners at 76 different sites. Each scanning session represented a different acquisition sequence of ≥2 back-to-back repeat scans (3.7 ± 0.7 on average). Unilateral HVs were determined with 18 different tools for automatic volumetry. An acquisition sequence was considered "poor" if the z-score of the within-session coefficient-of-variation of the HV estimates from the session, averaged across all volumetry tools and both hemispheres, exceeded one standard deviation. General IQM were computed for each scanning session using the freely available MRI Quality Control Tool. A classification-and-regression tree (CART) was trained to discriminate between good and poor acquisition sequences using the IQM as input.ResultsThe CART selected the left-right width of the acquisition field-of-view and the contrast-to-noise ratio as predictor variables. Overall accuracy of the CART was 79.5%. CART-based classification increased the ratio of good-to-poor acquisition sequences from 3.5 among all sequences to 7.4 among the sequences predicted to be good. This was at the expense of losing 15% of the good sequences.ConclusionsThe IQM-based decision tree model provides useful performance for the differentiation of T1-weighted sequences associated with good versus poor test-retest stability of hippocampus volumetry.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".