PD67 Strengthening And Accelerating Health Technology Assessments Through Artificial Intelligence
Bibliographic record
Abstract
Introduction: Rising costs and the rapidly increasing volume of findings from research in health care are driving the demand for comprehensive information to inform the allocation of resources. Health technology assessment (HTA) applies rigorous processes to provide high-quality synthesized information to policymakers and healthcare payers. HTA involves combining large amounts of research publications to systematically evaluate the properties, effects, and impacts on a topic of interest. Methods: The time and resources required to complete a full HTA are often demanding. There is an opportunity to apply high-performance computing (inclusive of artificial intelligence and machine learning disciplines) to HTA. This project applied high-computing technology to create a research synthesis tool to support HTA and then developed a service that integrates as much relevant data as possible to strengthen HTA. This was a joint project that combined expertise from the areas of health technology, machine learning, information technology, and innovation. Results: The information gathered for this phased project from HTA subject matter experts and other stakeholders was collated to inform a research synthesis tool and a broader concept of the project. Conclusions: The results of this study will inform the design of a research synthesis tool that covers the entire HTA process (literature search, screening titles and abstracts, data extraction, quality assessment, and analysis). The collaborators included Alberta Innovates, the Alberta Machine Intelligence Institute, the University of Alberta, Cybera, and PolicyWise. Alberta Innovates, which is an accelerator and innovator of research in the province of Alberta, Canada, was the primary source of funding for this project.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.241 | 0.259 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.014 | 0.013 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.016 | 0.007 |
| Open science | 0.004 | 0.014 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.021 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".