Health Condition Impacts in a Nationally Representative Cross-Sectional Survey Vary Substantially by Preference-Based Health Index
Bibliographic record
Abstract
IMPORTANCE: Many cost-utility analyses rely on generic utility measures for estimates of disease impact. Commonly used generic preference-based indexes may generate different absolute estimates of disease burden despite sharing anchors of dead at 0 and full health at 1.0. OBJECTIVE: We compare the impact of 16 prevalent chronic health conditions using 6 utility-based indexes of health and a visual analog scale. DESIGN: Data were from the National Health Measurement Study (NHMS), a cross-sectional telephone survey of 3844 adults aged 35 to 89 years in the United States. MAIN OUTCOME MEASURES: The NHMS included the EuroQol-5D-3L, Health and Activities Limitation Index (HALex), Health Utilities Index Mark 2 (HUI2) and Mark 3 (HUI3), preference-based scoring for the SF-36v2 (SF-6D), Quality of Well-Being Scale, and visual analog scale. Respondents self-reported 16 chronic conditions. Survey-weighted regression analyses for each index with all health conditions, age, and sex were used to estimate health condition impact estimates in terms of quality-adjusted life years (QALYs) lost over 10 years. All analyses were stratified by ages 35 to 69 and 70 to 89 years. RESULTS: There were significant differences between the indexes for estimates of the absolute impact of most conditions. On average, condition impacts were the smallest with the SF-6D and EQ-5D-3L and the largest with the HALex and HUI3. Likewise, the estimated loss of QALYs varied across indexes. Condition impact estimates for EQ-5D-3L, HUI2, HUI3, and SF-6D generally had strong Spearman correlations across conditions (i.e., >0.69). LIMITATIONS: This analysis uses cross-sectional data and lacks health condition severity information. CONCLUSIONS: Health condition impact estimates vary substantially across the indexes. These results imply that it is difficult to standardize results across cost-utility analyses that use different utility measures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.066 | 0.032 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".