Genome Diagnostics: Novel Strategies for Measuring Value
Bibliographic record
Abstract
Genetic testing technology is rapidly evolving with the growth of personalized medicine. While test evaluation typically relies on laboratory measures of performance, tests can be costly and analytically and ethically complex. A more fulsome consideration of value is warranted to inform adoption and appropriate use. Herein we describe a methodology for developing novel clinician- and patient-reported measures of clinical and personal utility, aiming to capture the informational value of genome diagnostic tests. Adhering to core measurement science principles and standards, our 4-step process includes (1) tool development through scoping reviews and stakeholder interviews and surveys; (2) tool validation through prospective cohort studies to establish construct validity, inter- and intra-rater reliability; (3) tool application using comparative effectiveness assessment to gauge the comparative value of different types of genetic tests; and (4) tool dissemination, leveraging existing partnerships with international stakeholders to spur additional validation studies, comparative effectiveness research, cost-effectiveness analysis, and evidence-informed policy. A scoping review of the clinical utility literature informed the development of a preliminary 25-item index. Qualitative interviews with 35 clinicians further informed the definition of our utility construct, item content, and item importance. Stakeholder surveys with 113 clinicians enabled further feedback on item content, importance, sensibility, response, and scoring options. An 18-item tool, the "Clinician-reported Genetic testing Utility InDEx" (C-GUIDE), is now undergoing validation, while development work on the patient-reported measure of utility is underway. A methodologically innovative approach to the development of stakeholder-informed and clinimetrically sound measures of value for personalized medicine tests will assist technology users and decision makers globally. DISCLOSURES: This work was supported by the Canadian Institutes of Health Research Operating Grant (#PJT-152880) and the PhRMA Foundation Challenge Award. Publication of the study methodology or findings generated therein was not contingent on the sponsor's approval or censorship of the manuscript. The authors have nothing to disclose. Results from this study were presented as a poster at the 40th Annual North American Meeting of the Society for Medical Decision Making; October 14, 2018; Montreal, QC; the Annual Meeting of the American Society of Human Genetics; October 18, 2018; San Diego, CA; and as an oral presentation at the Annual Meeting of the Canadian Association for Health Services and Policy Research; May 31, 2018; Montreal, QC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".