Using syndrome mining with the Health and Retirement Study to identify the deadliest and least deadly frailty syndromes
Bibliographic record
Abstract
Syndromes are defined with signs or symptoms that occur together and represent conditions. We use a data-driven approach to identify the deadliest and most death-averse frailty syndromes based on frailty symptoms. A list of 72 frailty symptoms was retrieved based on three frailty indices. We used data from the Health and Retirement Study (HRS), a longitudinal study following Americans aged 50 years and over. Principal component (PC)-based syndromes were derived based on a principal component analysis of the symptoms. Equal-weight 4-item syndromes were the sum of any four symptoms. Discrete-time survival analysis was conducted to compare the predictive power of derived syndromes on mortality. Deadly syndromes were those that significantly predicted mortality with positive regression coefficients and death-averse ones with negative coefficients. There were 2,797 of 5,041 PC-based and 964,774 of 971,635 equal-weight 4-item syndromes significantly associated with mortality. The input symptoms with the largest regression coefficients could be summed with three other input variables with small regression coefficients to constitute the leading deadliest and the most death-averse 4-item equal-weight syndromes. In addition to chance alone, input symptoms' variances and the regression coefficients or p values regarding mortality prediction are associated with the identification of significant syndromes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".