Automated fiber tract reconstruction for surgery planning: Extensive validation in language-related white matter tracts
Bibliographic record
Abstract
Diffusion MRI and tractography hold great potential for surgery planning, especially to preserve eloquent white matter during resections. However, fiber tract reconstruction requires an expert with detailed understanding of neuroanatomy. Several automated approaches have been proposed, using different strategies to reconstruct the white matter tracts in a supervised fashion. However, validation is often limited to comparison with manual delineation by overlap-based measures, which is limited in characterizing morphological and topological differences. In this work, we set up a fully automated pipeline based on anatomical criteria that does not require manual intervention, taking advantage of atlas-based criteria and advanced acquisition protocols available on clinical-grade MRI scanners. Then, we extensively validated it on epilepsy patients with specific focus on language-related bundles. The validation procedure encompasses different approaches, including simple overlap with manual segmentations from two experts, feasibility ratings from external multiple clinical raters and relation with task-based functional MRI. Overall, our results demonstrate good quantitative agreement between automated and manual segmentation, in most cases better performances of the proposed method in qualitative terms, and meaningful relationships with task-based fMRI. In addition, we observed significant differences between experts in terms of both manual segmentation and external ratings. These results offer important insights on how different levels of validation complement each other, supporting the idea that overlap-based measures, although quantitative, do not offer a full perspective on the similarities and differences between automated and manual methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".