Region-informed machine learning model for choroid plexus segmentation in Alzheimer’s disease
Bibliographic record
Abstract
Introduction The choroid plexus (CP), a critical structure for cerebrospinal fluid (CSF) production, has been increasingly recognized for its involvement in Alzheimer’s disease (AD). Accurate segmentation of CP from magnetic resonance imaging (MRI) remains challenging due to its irregular shape, variable MR signal, and proximity to the lateral ventricles. This study aimed to develop and evaluate a region-informed Gaussian Mixture Model (One-GMM) for automatic CP segmentation using anatomical priors derived from FreeSurfer (FS) software and compare it with manual, FS, and one previous GMM-based (Two-GMM) methods. Materials and methods T1-weighted (T1w) and T2-fluid-attenuated inversion recovery (FLAIR) MRI scans were acquired from 38 participants [19 cognitively normal (CN)], 11 with mild cognitive impairment (MCI), and 8 with AD. Manual segmentations served as ground truth. A GMM was applied within an anatomically constrained region combining the lateral ventricles and CP derived from FS reconstruction. The segmentation accuracy was assessed using the dice similarity coefficient (DSC), the 95th percentile Hausdorff distance (HD95), and volume difference percentage (VD%). Results were compared with those from FS and one previous GMM method-based segmentations across diagnostic groups. Results The region-informed One-GMM achieved significantly higher accuracy compared to FS and Two-GMM, with a mean DSC of 0.82 ± 0.05 for One-GMM versus 0.24 ± 0.11 for FS (p < 0.001), and 0.66 ± 0.10 for Two-GMM (p < 0.001), lower HD95 (One-GMM: 6.06 ± 10.32 mm vs. FS: 26.21 ± 7.32 mm vs. Two-GMM: 10.58 ± 6.47 mm), and comparable volume difference (One-GMM: 20.97 ± 9.53% vs. FS: 24.32 ± 28.13% vs. Two-GMM: 24.27 ± 22.10, p = 0.87). Segmentation accuracy of our proposed method was consistent across all diagnostic groups. Clinical analysis showed that there is no diagnostic group difference in CP volume obtained from manual, FS, Two-GMM, and our proposed One-GMM methods. In the whole cohort, there are also no age and sex effects of CP volume with all methods. Restricting to the CN group, CP volume from both manual (p = 0.03), Two-GMM (p < 0.01) and the proposed One-GMM (p = 0.05), methods show an aging effect, but not for the FS segmented CP volume (p = 0.22). Conclusion A region-informed One-GMM method significantly improved CP segmentation accuracy over FS, providing a practical and accessible tool for CP quantification in AD and other research studies. Within this small cohort, no diagnostic group difference in CP volume was observed. An aging effect of CP volume was found within the CN group.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".