Appraisals by Health Technology Assessment Agencies of Economic Evaluations Submitted as Part of Reimbursement Dossiers for Oncology Treatments: Evidence from Canada, the UK, and Australia
Bibliographic record
Abstract
Publicly funded healthcare systems, including those in Canada, the United Kingdom (UK), and Australia, often use health technology assessment (HTA) to inform drug reimbursement decision-making, based on dossiers submitted by manufacturers, and HTA agencies issue publicly available reports to support funding recommendations. However, the level of information reported by HTA agencies in these reports may vary. To provide insights on this issue, we describe and assess the reporting of economic methods in recent oncology HTA recommendations from the Canadian Agency for Drugs and Technologies in Health (CADTH), National Institute for Health and Care Excellence (NICE), and Pharmaceutical Benefits Advisory Committee (PBAC). Publicly available HTA recommendations and reports for oncology drugs issued by CADTH over a 2-year period, 2019-2020, were identified and compared with the corresponding HTA documents from NICE and the PBAC. Reporting of key model characteristics and attributes, survival analysis methods, methodological criticisms, and re-assessment of the economic results were characterized using descriptive statistics. Dichotomous differences in the methodological criticisms observed between the three agencies were assessed using Cochran's Q tests and substantiated using pairwise McNemar tests. Chi-squared tests were used to assess the dichotomous differences in the reporting of methods and explore the potential relationships between categorical variables, where appropriate. HTAs published by CADTH, NICE, and the PBAC consistently reported a broad spectrum of descriptive information on the economic models submitted by manufacturers. While common economic evaluation attributes were well-reported across the three HTA agencies, significant differences in the reporting of survival analysis methods and methodological criticisms were observed. NICE consistently reported more comprehensive information, compared to either CADTH or PBAC. Despite these differences, broadly similar recommendation rates were observed between CADTH and NICE. The PBAC was found to be more restrictive. Based on our 2-year sample of oncology, the HTAs published by CADTH matched with the corresponding HTAs from NICE and PBAC; we observed important variations in the reporting of economic evidence, especially technical aspects, such as survival analysis, across the three agencies. In addition to guidelines for HTA submissions by manufacturers, the community of HTA agencies should also have common standards for reporting the results of their assessments, though the information and opinions reported may differ.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.626 | 0.898 |
| Meta-epidemiology (narrow) | 0.002 | 0.004 |
| Meta-epidemiology (broad) | 0.007 | 0.018 |
| Bibliometrics | 0.039 | 0.048 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.025 | 0.010 |
| Open science | 0.007 | 0.010 |
| Research integrity | 0.005 | 0.010 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".