An objective metric of individual health and aging for population surveys
Bibliographic record
Abstract
BACKGROUND: We have previously developed and validated a biomarker-based metric of overall health status using Mahalanobis distance (DM) to measure how far from the norm of a reference population (RP) an individual's biomarker profile is. DM is not particularly sensitive to the choice of biomarkers; however, this makes comparison across studies difficult. Here we aimed to identify and validate a standard, optimized version of DM that would be highly stable across populations, while using fewer and more commonly measured biomarkers. METHODS: Using three datasets (the Baltimore Longitudinal Study of Aging, Invecchiare in Chianti and the National Health and Nutrition Examination Survey), we selected the most stable sets of biomarkers in all three populations, notably when interchanging RPs across populations. We performed regression models, using a fourth dataset (the Women's Health and Aging Study), to compare the new DM sets to other well-known metrics [allostatic load (AL) and self-assessed health (SAH)] in their association with diverse health outcomes: mortality, frailty, cardiovascular disease (CVD), diabetes, and comorbidity number. RESULTS: A nine- (DM9) and a seventeen-biomarker set (DM17) were identified as highly stable regardless of the chosen RP (e.g.: mean correlation among versions generated by interchanging RPs across dataset of r = 0.94 for both DM9 and DM17). In general, DM17 and DM9 were both competitive compared with AL and SAH in predicting aging correlates, with some exceptions for DM9. For example, DM9, DM17, AL, and SAH all predicted mortality to a similar extent (ranges of hazard ratios of 1.15-1.30, 1.21-1.36, 1.17-1.38, and 1.17-1.49, respectively). On the other hand, DM9 predicted CVD less well than DM17 (ranges of odds ratios of 0.97-1.08, 1.07-1.85, respectively). CONCLUSIONS: The metrics we propose here are easy to measure with data that are already available in a wide array of panel, cohort, and clinical studies. The standardized versions here lose a small amount of predictive power compared to more complete versions, but are nonetheless competitive with existing metrics of overall health. DM17 performs slightly better than DM9 and should be preferred in most cases, but DM9 may still be used when a more limited number of biomarkers is available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.059 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".