Cost-effectiveness thresholds in health care: a bookshelf guide to their meaning and use
Bibliographic record
Abstract
There is misunderstanding about both the meaning and the role of cost-effectiveness thresholds in policy decision making. This article dissects the main issues by use of a bookshelf metaphor. Its main conclusions are as follows: it must be possible to compare interventions in terms of their impact on a common measure of health; mere effectiveness is not a persuasive case for inclusion in public insurance plans; public health advocates need to address issues of relative effectiveness; a 'first best' benchmark or threshold ratio of health gain to expenditure identifies the least effective intervention that should be included in a public insurance plan; the reciprocal of this ratio - the 'first best' cost-effectiveness threshold - will rise or fall as the health budget rises or falls (ceteris paribus); setting thresholds too high or too low costs lives; failure to set any cost-effectiveness threshold at all also involves avertable deaths and morbidity; the threshold cannot be set independently of the health budget; the threshold can be approached from either the demand side or the supply side - the two are equivalent only in a health-maximising equilibrium; the supply-side approach generates an estimate of a 'second best' cost-effectiveness threshold that is higher than the 'first best'; the second best threshold is the one generally to be preferred in decisions about adding or subtracting interventions in an established public insurance package; multiple thresholds are implied by systems having distinct and separable health budgets; disinvestment involves eliminating effective technologies from the insured bundle; differential weighting of beneficiaries' health gains may affect the threshold; anonymity and identity are factors that may affect the interpretation of the threshold; the true opportunity cost of health care in a community, where the effectiveness of interventions is determined by their impact on health, is not to be measured in money - but in health itself.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".