DEVELOPING AND TESTING A TOOL FOR ASSESSING PHARMACOECONOMIC EVALUATIONS
Bibliographic record
Abstract
Introduction Manufacturers seeking reimbursement from publicly funded drug plans submit a drug submission to the Common Drug Review (CDR). This includes clinical data and a pharmacoeconomic evaluation (PE) which is assessed by clinical and economic teams. Clinical and pharmacoeconomic reports are used by the Canadian Drug Expert Committee (CDEC) in their deliberations, which could result in a “do not list”, “list with criteria” or “list” formulary listing recommendation (FLR). Given growing financial constraints, the need to economic information to inform FLR decision makers is critical. Therefore, understanding the quality and limitations with regards to economic submissions can assist FLR decision makers. This report focuses on the development and piloting of a PE appraisal instrument. Objectives To develop an instrument for assessing the quality of PEs in terms of adherence to guidelines, methodological quality and transparency of reporting, and use of indirect treatment comparisons. Methods An existing PE appraisal instrument was updated using a quality assessment framework (QAF). A literature review was conducted to identify additional key factors affecting methodological quality. Modifications to the tool were made upon consensus. The instrument was piloted by three investigators using 29 selected PEs and accompanying pharmacoeconomic reports. A pooled kappa statistic was calculated. The instrument was evaluated by identifying items met from QAF and comparing it to four highly-recognised appraisal tools. Results This tool includes 69 multiple-choice and short answer questions and covers: economic evaluation details, quality of pharmacoeconomic submissions, and use of pharmacoeconomic information by CDEC. The kappa statistic indicates substantial agreement [0.72(0.67–0.77)]. The tool fulfils all dimensions of QAF and is at least as comprehensive as other published appraisal instruments. Conclusions The steps in developing this PE quality appraisal instrument involved: quality assessment framework, literature review, draft tool, pretesting, re-drafted instrument, inter-rater reliability testing, and final instrument. Although this tool was created to examine PEs submitted to CDR, components of this instrument can be used beyond the realm of CDR PE assessments. Its practicality, reliability and parity to other appraisal tools demonstrate its potential application in the development and appraisal of PEs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.341 | 0.545 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.007 |
| Bibliometrics | 0.022 | 0.016 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.008 | 0.010 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".