Artificial intelligence for automatic cerebral ventricle segmentation and volume calculation: a clinical tool for the evaluation of pediatric hydrocephalus
Bibliographic record
Abstract
OBJECTIVE: Imaging evaluation of the cerebral ventricles is important for clinical decision-making in pediatric hydrocephalus. Although quantitative measurements of ventricular size, over time, can facilitate objective comparison, automated tools for calculating ventricular volume are not structured for clinical use. The authors aimed to develop a fully automated deep learning (DL) model for pediatric cerebral ventricle segmentation and volume calculation for widespread clinical implementation across multiple hospitals. METHODS: The study cohort consisted of 200 children with obstructive hydrocephalus from four pediatric hospitals, along with 199 controls. Manual ventricle segmentation and volume calculation values served as "ground truth" data. An encoder-decoder convolutional neural network architecture, in which T2-weighted MR images were used as input, automatically delineated the ventricles and output volumetric measurements. On a held-out test set, segmentation accuracy was assessed using the Dice similarity coefficient (0 to 1) and volume calculation was assessed using linear regression. Model generalizability was evaluated on an external MRI data set from a fifth hospital. The DL model performance was compared against FreeSurfer research segmentation software. RESULTS: Model segmentation performed with an overall Dice score of 0.901 (0.946 in hydrocephalus, 0.856 in controls). The model generalized to external MR images from a fifth pediatric hospital with a Dice score of 0.926. The model was more accurate than FreeSurfer, with faster operating times (1.48 seconds per scan). CONCLUSIONS: The authors present a DL model for automatic ventricle segmentation and volume calculation that is more accurate and rapid than currently available methods. With near-immediate volumetric output and reliable performance across institutional scanner types, this model can be adapted to the real-time clinical evaluation of hydrocephalus and improve clinician workflow.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".