Are we properly evaluating genetic and genomic testing? A systematic review of health technology assessment reports
Bibliographic record
Abstract
BACKGROUND: Despite advances in precision medicine, the translation of genetic and genomic technologies into routine practice is hampered by a heterogeneous and limited evidence base and the absence of standardized evaluation methodologies. Health Technology Assessment (HTA) plays a critical role in bridging this gap, yet assessment approaches and comprehensiveness vary widely. This systematic review aims to map the landscape of the assessment reports on genetic and genomics applications, analyze their methodological aspects and identify gaps. METHODS: PubMed, Scopus, Web of Science, and the international HTA database, were searched for assessment reports of genetic/genomic technologies. Information on reports general characteristics, assessment domains and their components, consulted sources of evidence and reported gaps was extracted. Findings were synthesized narratively. RESULTS: Out of 27,331 screened records, 41 reports were included, predominantly from Canada, the United Kingdom, and Australia, mainly aimed at informing policy making for single or multiple gene tests for cancer patients. Most reports used a generic HTA methodology and assessment domains varied across reports. Key clinical aspects, such as clinical accuracy and safety, suffered from evidence gaps (39.0% and 22.0%), while personal and societal aspects were the least investigated assessment domain (48.8-78.0%). Overall, lack of evidence and limited generalizability of findings were the most commonly reported gaps across multiple domains. CONCLUSIONS: The review highlighted significant fragmentation in current evaluation methodologies of genetic and genomic applications, with underassessment of analytical/clinical accuracy, safety, and non-health outcomes, alongside evidence gaps and limited generalizability. These issues compromise both evaluation and decision-making process, underscoring the urgent need for alternative study designs and standardized, comprehensive assessment frameworks to facilitate the successful implementation of emerging genetic and genomic technologies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".