Three Challenges to Achieving Better Analysis for Better Decisions: Generalisability, Complexity and Thresholds
Bibliographic record
Abstract
In June 2006, a conference entitled Better Analysis for Better Decisions - Bridging the Gap Between Economic Evaluation and Healthcare Decision-Making was held at McMaster University in honour of the late Bernie O'Brien. The papers presented by leading health economists were reviews of the use of economic evaluation in the UK, Canada and USA, and more methodologically focused contributions. The reviews of the experience in the three countries suggest that economic analysis is playing an increasingly important role in health sector decision making. But usage is patchy, rarely systematic and explicit, reflecting both some irrational illogical resistance and some genuine concerns and problems with the current methods of analysis. We identified three key issues1 that recurred in the papers and in the discussion at the conference - • generalisability - the extent to which cost-effectiveness analyses relevant and appropriate to one jurisdiction can be used in another; • complexity - two related issues arise as economic evaluations become more complex - - credibility with decision makers and - the need to trade complexity against quantity to enable the analyses of more technologies within Health Technology Assessment (HTA) resource constraints; • thresholds - the basis for, and validity of, thresholds values for the incremental cost-effectiveness ratio (e.g. cost per Quality Adjusted Life Year (QALY)) adopted by central decision-makers and their relevance at a local level. The challenges posed by these issues result from success, i.e. the greater use of economic evaluation in decision making. Meeting them is fundamental to its continued growth in use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.827 | 0.863 |
| Meta-epidemiology (narrow) | 0.006 | 0.005 |
| Meta-epidemiology (broad) | 0.020 | 0.009 |
| Bibliometrics | 0.016 | 0.012 |
| Science and technology studies | 0.008 | 0.093 |
| Scholarly communication | 0.038 | 0.074 |
| Open science | 0.017 | 0.037 |
| Research integrity | 0.028 | 0.062 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".