From a genomic risk model to clinical trial implementation in a learning health system: the ProGRESS Study
Bibliographic record
Abstract
ABSTRACT Background As healthcare moves from a one-size-fits-all approach towards precision care, individual risk prediction is an important step in disease prevention and early detection. Biobank-linked healthcare systems can generate knowledge about genomic risk and test the impact of implementing that knowledge in care. Risk-stratified prostate cancer screening is one clinical application that might benefit from such an approach. Methods We developed a clinical translation pipeline for genomics-informed prostate cancer screening in a national healthcare system. We used data from 585,418 male participants of the Veterans Affairs (VA) Million Veteran Program (MVP), among whom 101,920 self-identify as Black/African-American, to develop and validate the Prostate CAncer integrated Risk Evaluation (P-CARE) model, a prostate cancer risk prediction model based on a polygenic score, family history, and genetic principal components. The model was externally validated in data from 18,457 PRACTICAL Consortium participants. A novel blended genome-exome (BGE) platform was used to develop a clinical laboratory assay for both the P-CARE model and rare variants in prostate cancer-associated genes, including additional validation in 74,331 samples from the All of Us Research Program. Results In overall and ancestry-stratified analyses, the polygenic score of 601 variants was associated with any, metastatic, and fatal prostate cancer in MVP and PRACTICAL. Values of the P-CARE model at ≥80th percentile in the multiancestry cohort overall were associated with hazard ratios (HR) of 2.75 (95% CI 2.66-2.84), 2.78 (95% CI 2.54-2.99), and 2.59 (95% CI 2.22-2.97) for any, metastatic, and fatal prostate cancer in MVP, respectively, compared to the median. When high– and low-risk groups were defined as P-CARE HR>1.5 and HR<0.75 for metastatic prostate cancer, the 220,062 (37.6%) high-risk vs.146,826 (25.1%) low-risk participants in MVP had a 47.9% vs. 14.1%, 9.3% vs. 2.0%, and 3.6% vs. 0.8% cumulative cause-specific incidence of any, metastatic, and fatal prostate cancer by age 90, respectively. The clinical assay and reports are now being implemented in a clinical trial of precision prostate cancer screening in the VA healthcare system (Clinicaltrials.gov NCT05926102 ). Conclusions A model consisting of a polygenic score, family history, and genetic principal components describes a clinically important gradient of prostate cancer risk in a diverse patient population and demonstrates the potential of learning health systems to implement and evaluate precision health care approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.131 | 0.310 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".