Identifying temporal patterns in the progression of neurodegenerative disease using unsupervised clustering
Bibliographic record
Abstract
Abstract Objectives One of the principal goals of Precision Medicine is to stratify patients by accounting for individual variability. However, extracting meaningful information from Real-World Data, such as Electronic Health Records, still remains challenging due to methodological and computational issues. Materials & methods A Dynamic Time Warping-based unsupervised-clustering methodology is presented in this paper for the clustering of patient trajectories of multi-modal health data on the basis of shared temporal characteristics. Building on an earlier methodology, a new dimension of time-varying numerical clinical and imaging features (six in total) is incorporated, through an adapted cost-minimization algorithm for clustering on different, possibly overlapping, feature subsets. A cluster evaluation process is also implemented, by admitting two user-defined parameters (granularity threshold and feature contribution). The model disease chosen is Huntington’s disease (HD), characterized by progressive neurodegeneration. Results From a wide range of examined user-defined parameters, four case examples are highlighted to exemplify the combined effects of feature weights and granularity threshold in the stratification of HD trajectories in homogeneous clusters. For each identified cluster, polynomial fits that describe the temporal behavior of the assessed features are provided for an informative comparison, together with their averaged values. Discussion The proposed data-mining methodology permits the stratification of distinct time patterns of multi-modal health data in individuals that share a diagnosis or future diagnosis, employing user-customized criteria beyond the current clinical practice. Conclusions This work bears implications for better analysis of individual variability in disease progression, opening doors to personalized preventative, diagnostic and therapeutic strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".