Parkinson’s Disease Subtypes: Critical Appraisal and Recommendations
Bibliographic record
Abstract
BACKGROUND: In Parkinson's disease (PD), there is heterogeneity in the clinical presentation and underlying biology. Research on PD subtypes aims to understand this heterogeneity with potential contribution for the knowledge of disease pathophysiology, natural history and therapeutic development. There have been many studies of PD subtypes but their impact remains unclear with limited application in research or clinical practice. OBJECTIVE: To critically evaluate PD subtyping systems. METHODS: We conducted a systematic review of PD subtypes, assessing the characteristics of the studies reporting a subtyping system for the first time. We completed a critical appraisal of their methodologic quality and clinical applicability using standardized checklists. RESULTS: We included 38 studies. The majority were cross-sectional (n = 26, 68.4%), used a data-driven approach (n = 25, 65.8%), and non-clinical biomarkers were rarely used (n = 5, 13.1%). Motor characteristics were the domain most commonly reported to differentiate PD subtypes. Most of the studies did not achieve the top rating across items of a Methodologic Quality checklist. In a Clinical Applicability Checklist, the clinical importance of differences between subtypes, potential treatment implications and applicability to the general population were rated poorly, and subtype stability over time and prognostic value were largely unknown. CONCLUSION: Subtyping studies undertaken to date have significant methodologic shortcomings and most have questionable clinical applicability and unknown biological relevance. The clinical and biological signature of PD may be unique to the individual, rendering PD resistant to meaningful cluster solutions. New approaches that acknowledge the individual-level heterogeneity and that are more aligned with personalized medicine are needed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".