Reevaluating health metrics: Unraveling the limitations of disability-adjusted life years as an indicator in disease burden assessment
Bibliographic record
Abstract
In 1993, the World Bank released a global report on the efficacy of health promotion, introducing the disability-adjusted life years (DALY) as a novel indicator. The DALY, a composite metric incorporating temporal and qualitative data, is grounded in preferences regarding disability status. This review delineates the algorithm used to calculate the value of the proposed DALY synthetic indicator and elucidates key methodological challenges associated with its application. In contrast to the quality-adjusted life years approach, derived from multi-attribute utility theory, the DALY stands as an independent synthetic indicator that adopts the assumptions of the Time Trade Off utility technique to define Disability Weights. Claiming to rely on no mathematical or economic theory, DALY users appear to have exempted themselves from verifying whether this indicator meets the classical properties required of all indicators, notably content validity, reliability, specificity, and sensitivity. The DALY concept emerged primarily to facilitate comparisons of the health impacts of various diseases globally within the framework of the Global Burden of Disease initiative, leading to numerous publications in international literature. Despite widespread adoption, the DALY synthetic indicator has prompted significant methodological concerns since its inception, manifesting in inconsistent and non-reproducible results. Given the substantial diffusion of the DALY indicator and its critical role in health impact assessments, a reassessment is warranted. This reconsideration is imperative for enhancing the robustness and reliability of public health decision-making processes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.160 | 0.056 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".