Using Effectiveness and Cost-effectiveness to Make Drug Coverage Decisions
Bibliographic record
Abstract
CONTEXT: National public insurance for drugs is often based on evidence of comparative effectiveness and cost-effectiveness. This study describes how that evidence has been used across 3 jurisdictions (Australia, Canada, and Britain) that have been at the forefront of evidence-based coverage internationally. OBJECTIVES: To describe how clinical and cost-effectiveness evidence is used in coverage decisions both within and across jurisdictions and to identify common issues in the process of evidence-based coverage. DESIGN, SETTING, AND PARTICIPANTS: Descriptive analysis of retrospective data from the Common Drug Review (CDR) of Canada, National Institute for Health and Clinical Excellence (NICE) in Britain, and Pharmaceutical Benefits Advisory Committee (PBAC) of Australia. All publicly available information as of December 31, 2008, was gathered from each committee's Web site (data set begins in January 2004 [CDR], February 2001 [NICE], and July 2005 [PBAC]). MAIN OUTCOME MEASURE: Listing recommendations for each drug by disease indication. RESULTS: NICE recommended 87.4% (174/199) of submissions for listing compared with a listing rate of 49.6% (60/121) and 54.3% (153/282) for the CDR and PBAC, respectively. Significant uncertainty around clinical effectiveness, typically resulting from inadequate study design or the use of inappropriate comparators and unvalidated surrogate end points, was identified as a key issue in coverage decisions. Recommendations varied considerably across countries, possibly because of differences in the medications reviewed; different agency processes, including the willingness to negotiate on price; and the approach to "me too" drugs. The data suggest that the 3 agencies make recommendations that are consistent with evidence on effectiveness and cost-effectiveness but that other factors are often important. CONCLUSIONS: NICE, PBAC, and CDR face common issues with respect to the quality and strength of the experimental evidence in support of a clinically meaningful effect. However, comparative effectiveness and cost-effectiveness, along with other relevant factors, can be used by national agencies to support drug decision making. The results of the evaluation process in different countries are influenced by the context, agency processes, ability to engage in price negotiation, and perhaps differences in social values.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".