Machine Learning for Subtyping Concussion Using a Clustering Approach
Bibliographic record
Abstract
Background: Concussion subtypes are typically organized into commonly affected symptom areas or a combination of affected systems, an approach that may be flawed by bias in conceptualization or the inherent limitations of interdisciplinary expertise. Objective: The purpose of this study was to determine whether a bottom-up, unsupervised, machine learning approach, could more accurately support concussion subtyping. Methods: Initial patient intake data as well as objective outcome measures including, the Patient-Reported Outcomes Measurement Information System (PROMIS), Dizziness Handicap Inventory (DHI), Pain Catastrophizing Scale (PCS), and Immediate Post-Concussion Assessment and Cognitive Testing Tool (ImPACT) were retrospectively extracted from the Advance Concussion Clinic's database. A correlation matrix and principal component analysis (PCA) were used to reduce the dimensionality of the dataset. Sklearn's agglomerative clustering algorithm was then applied, and the optimal number of clusters within the patient database were generated. Between-group comparisons among the formed clusters were performed using a Mann-Whitney U test. Results: Two hundred seventy-five patients within the clinics database were analyzed. Five distinct clusters emerged from the data when maximizing the Silhouette score (0.36) and minimizing the Davies-Bouldin score (0.83). Concussion subtypes derived demonstrated clinically distinct profiles, with statistically significant differences ( p < 0.05) between all five clusters. Conclusion: This machine learning approach enabled the identification and characterization of five distinct concussion subtypes, which were best understood according to levels of complexity, ranging from Extremely Complex to Minimally Complex. Understanding concussion in terms of Complexity with the utilization of artificial intelligence, could provide a more accurate concussion classification or subtype approach; one that better reflects the true heterogeneity and complex system disruptions associated with mild traumatic brain injury.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".