Merged Group Tractography Evaluation with Selective Automated Group Integrated Tractography
Bibliographic record
Abstract
Introduction: Tractography analysis in group-based studies across large populations is difficult to manage and assess. We propose Selective Automated Group Integrated Tractography (SAGIT), an automated group tractography software platform that incorporates proven dMRI practices, in order for group-wise dMRI to be more accessible to researchers. We use a merged tractography approach that permits evaluation of tractography datasets at the group level. We also introduce an image-based score (Normalized Overlapping Score (NOS)) that can quantify the quality of the group tractography results. We deploy SAGIT to evaluate deterministic and probabilistic constrained spherical deconvolution (CSTdet, CSTprob), extended streamline tractography (XST), and diffusion tensor tractography (DTT) in their ability to delineate different neuroanatomy, as well as validating NOS across these different brain regions. Methods: MR sequences were acquired from 42 healthy adults. Anatomical and group registrations were performed using ANTs. Cortical segmentation was performed using FreeSurfer. Four tractography algorithms were used to delineate 6 sets of neuroanatomy: fornix, facial/vestibular-cochlear cranial nerve complex, vagus nerve, rubral-cerebellar decussation, optic radiation, and auditory radiation. The tracts were generated both with and without ROI filters. The generated visual reports were then evaluated by 5 neuroscientists. Results: At a group level, merged tractography demonstrated that different methods have different fiber distribution characteristics. CSTprob is prone to false-positives, and thereby suitable in anatomy with strong priors. CSTdet and XST are more conservative, but have more trouble resolving hemispherical decussation and distant crossing projections. DTT consistently shows the worst reproducibility across the anatomies. Linear regression of rater scores against NOS shows significant (p0.05) for unfiltered tractography. Conclusions: The tractography results demonstrated reliable and consistent performance of SAGIT across multiple subjects and techniques. Through SAGIT, we quantifiably demonstrated that different algorithms showed different strength and weaknesses at a group level. No single algorithm seems to be suitable for all anatomical tasks, it would be useful to consider using a mix of algorithms for different anatomical segments. Finally, we have demonstrated that merged tractography is a promising group-wise tractography analysis approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".