131 Toward a living model for health technology assessments
Bibliographic record
Abstract
<h3>Objectives</h3> Health care policy should not be based on outdated evidence. Keeping people healthy requires up-to-date evidence to inform coverage decisions. Health technology assessments (HTAs) drive care pathways by using systematic review (SR) methodology to determine payment policy for clinical interventions using public input. HTA reports are updated on an as-needed basis, similar to SRs and clinical practice guidelines. However, as the rate of health-related publications has increased, so too has the need to keep evidence syntheses up to date using the latest evidence on a more frequent, ‘living’ basis. We seek to create a model for a living HTA process. <h3>Method</h3> As HTAs are not always published in journals, we searched grey literature but found only two living HTAs, from Canada, that used Cochrane’s guidance for living systematic reviews (LSRs). Compared with other published guidance, the Cochrane guidance captured all current best practices for LSRs. We used the Cochrane guidance as the basis for the living HTA model. The three core tenets of LSRs include: regularly monitoring the evidence base; incorporating new evidence on a pre-determined threshold; and transparently communicating update status. It is an approach to updating reviews, not a review type or method, which can be translated directly to inform living HTAs, meaning reviews may transition in and out of a living state. <h3>Results</h3> Adapting the guidance for LSRs to living HTAs was uncomplicated due to the overlap of SRs and HTAs. Cochrane states that a living model is best when the review question is a particular priority for decision-making, there is an important level of uncertainty in existing evidence, and there will likely be emerging evidence to impact the decision. The first is true for all HTA topics, as they are selected for their decision-making priority. A potential hindrance is the update process itself; for example, HTA topics chosen for re-review by Washington State must proceed through a formal process as outlined in state law. Incorporating new evidence into an HTA can be suggested, but the decision does not lie solely with the HTA program; re-reviewing a topic can be a drawn-out process. This differs from updating an LSR, which can begin as soon as new evidence is identified. <h3>Conclusions</h3> Creating guidance for living HTAs led to an objective method for determining re-reviews, a transparent state of communicating HTA updates, and a consistent workflow for those involved in the HTA process. As topics are prioritized to transition to a living state, the efficacy and cost-effectiveness of this approach can be assessed. Clear criteria and specific triggers of an HTA update may reduce overall workload and costs but will be determined using longitudinal data. Usefulness of the model, as well as individual thresholds for updates, will need to be assessed annually; legal requirements may hinder the speed of decision-making. As more HTA topics transition into or out of a living state, criteria for updating or ceasing updates can be further refined, along with customization of search frequency for each topic.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".