Exclusion Criteria in National Health State Valuation Studies
Bibliographic record
Abstract
BACKGROUND: Health state valuation data are often excluded from studies that aim to provide a nationally representative set of values for preference-based health-related quality of life (HRQoL) instruments. The purpose was to provide a systematic examination of exclusion criteria used in the derivation of societal scoring algorithms for preference-based HRQoL instruments. METHODS: Data sources included MEDLINE, official instrument websites, and publication reference lists. Analyses that used data from national valuation studies and reported a scoring algorithm for a generic preference-based HRQoL instrument were included. Data extraction included exclusion criteria and associated justifications, exclusion rates, the characteristics of excluded respondents, and analyses that explored consequential implications of exclusion criteria on the respective national tariff. RESULTS: Seventy-six analyses (from 70 papers) met the inclusion criteria. In addition to being excluded for logical inconsistencies, respondents were often excluded if they valued fewer than 3 health states or if they gave the same value to all health states. Numerous other exclusion criteria were identified, with varying degrees of justification, often based on an assumption that respondents did not understand the task or as a consequence of the chosen statistical modeling techniques. Rates of exclusion ranged from 0% to 65%, with excluded respondents more likely to be older, less educated, and less healthy. Limitations included that the database search was confined to MEDLINE; study selection focused on national valuation studies that used standard gamble, time tradeoff, and/or visual analog scale techniques; and only English-language studies were included. CONCLUSION: Exclusion criteria used in national valuation studies vary considerably. Further consideration is necessary in this important and influential area of research, from the design stage to the reporting of results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.116 | 0.061 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".