NigraNet: An automatic framework to assess nigral neuromelanin content in early Parkinson’s disease using convolutional neural network
Bibliographic record
Abstract
Parkinson’s disease (PD) demonstrates neurodegenerative changes in the substantia nigra pars compacta (SNc) using neuromelanin-sensitive (NM)-MRI. As SNc manual segmentation is prone to substantial inter-individual variability across raters, development of a robust automatic segmentation framework is necessary to facilitate nigral neuromelanin quantification. Artificial intelligence (AI) is gaining traction in the neuroimaging community for automated brain region segmentation tasks using MRI. Developing and validating AI-based NigraNet, a fully automatic SNc segmentation framework allowing nigral neuromelanin quantification in patients with PD using NM-MRI. We prospectively included 199 participants comprising 144 early-stage idiopathic PD patients (disease duration = 1.5 ± 1.0 years) and 55 healthy volunteers (HV) scanned using a 3 Tesla MRI including whole brain T1-weighted anatomical imaging and NM-MRI. The regions of interest (ROI) were delineated in all participants automatically using NigraNet, a modified U-net, and compared to manual segmentations performed by two experienced raters. The SNc volumes (Vol), volumes corrected by total intracranial volume (Cvol), normalized signal intensity (NSI) and contrast-to-noise ratio (CNR) were computed. One-way GLM-ANCOVA was performed while adjusting for age and sex as covariates. Diagnostic performance measurement was assessed using the receiver operating characteristic (ROC) analysis. Inter and intra-observer variability were estimated using Dice similarity coefficient (DSC). The agreements between methods were tested using intraclass correlation coefficient (ICC) based on a mean-rating, two-way, mixed-effects model estimates for absolute agreement. Cronbach’s alpha and Bland-Altman plots were estimated to assess inter-method consistency. Using both methods, Vol, Cvol, NSI and CNR measurements differed between PD and HV with an effect of sex for Cvol and CNR. ICC values between the methods demonstrated optimal agreement for Cvol and CNR (ICC > 0.9) and high reproducibility (DSC: 0.80) was also obtained. The SNc measurements also showed good to excellent consistency values (Cronbach's alpha > 0.87). Bland-Altman plots of agreement demonstrated no association of SNc ROI measurement differences between the methods and ROI average measurements while confirming that 95% of the data points were ranging between the limits of mean difference (d ± 1.96xSD). Percentage changes between PD and HV were -27.4% and -17.7% for Vol, -30.0% and -22.2% for Cvol, -15.8% and -14.4% for NSI, -17.1% and -16.0% for CNR for automatic and manual measurements respectively. Using automatic method, in the entire dataset, we obtained the areas under the ROC curve (AUC) of 0.83 for Vol, 0.85 for Cvol, 0.79 for NSI and 0.77 for CNR whereas in the training dataset of 0.96 for Vol, 0.95 for Cvol, 0.85 for NSI and 0.85 for CNR. Disease duration correlated negatively with NSI of the patients for both the automatic and manual measurements. We presented an AI-based NigraNet framework that utilizes a small MRI training dataset to fully automatize the SNc segmentation procedure with an increased precision and more reproducible results. Considering the consistency, accuracy and speed of our approach, this study could be a crucial step towards the implementation of a time-saving non-rater dependent fully automatic method for studying neuromelanin changes in clinical settings and large-scale neuroimaging studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".