African-specific improvement of a polygenic hazard score for age at diagnosis of prostate cancer
Bibliographic record
Abstract
Polygenic hazard score (PHS) models are associated with age at diagnosis of prostate cancer. Our model developed in Europeans (PHS46) showed reduced performance in men with African genetic ancestry. We used a cross-validated search to identify single nucleotide polymorphisms (SNPs) that might improve performance in this population. Anonymized genotypic data were obtained from the PRACTICAL consortium for 6253 men with African genetic ancestry. Ten iterations of a 10-fold cross-validation search were conducted to select SNPs that would be included in the final PHS46+African model. The coefficients of PHS46+African were estimated in a Cox proportional hazards framework using age at diagnosis as the dependent variable and PHS46, and selected SNPs as predictors. The performance of PHS46 and PHS46+African was compared using the same cross-validated approach. Three SNPs (rs76229939, rs74421890 and rs5013678) were selected for inclusion in PHS46+African. All three SNPs are located on chromosome 8q24. PHS46+African showed substantial improvements in all performance metrics measured, including a 75% increase in the relative hazard of those in the upper 20% compared to the bottom 20% (2.47-4.34) and a 20% reduction in the relative hazard of those in the bottom 20% compared to the middle 40% (0.65-0.53). In conclusion, we identified three SNPs that substantially improved the association of PHS46 with age at diagnosis of prostate cancer in men with African genetic ancestry to levels comparable to Europeans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".