A risk prediction model for head and neck cancers incorporating lifestyle factors, <scp>HPV</scp> serology and genetic markers
Bibliographic record
Abstract
Head and neck cancer is often diagnosed late and prognosis for most head and neck cancer patients remains poor. To aid early detection, we developed a risk prediction model based on demographic and lifestyle risk factors, human papillomavirus (HPV) serological markers and genetic markers. A total of 10 126 head and neck cancer cases and 5254 controls from five North American and European studies were included. HPV serostatus was determined by antibodies for HPV16 early oncoproteins (E6, E7) and regulatory early proteins (E1, E2, E4). The data were split into a training set (70%) for model development and a hold-out testing set (30%) for model performance evaluation, including discriminative ability and calibration. The risk models including demographic, lifestyle risk factors and polygenic risk score showed a reasonable predictive accuracy for head and neck cancer overall. A risk model that also included HPV serology showed substantially improved predictive accuracy for oropharyngeal cancer (AUC = 0.94, 95% CI = 0.92-0.95 in men and AUC = 0.92, 95% CI = 0.88-0.95 in women). The 5-year absolute risk estimates showed distinct trajectories by risk factor profiles. Based on the UK Biobank cohort, the risks of developing oropharyngeal cancer among 60 years old and HPV16 seropositive in the next 5 years ranged from 5.8% to 14.9% with an average of 8.1% for men, 1.3% to 4.4% with an average of 2.2% for women. Absolute risk was generally higher among individuals with heavy smoking, heavy drinking, HPV seropositivity and those with higher polygenic risk score. These risk models may be helpful for identifying people at high risk of developing head and neck cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".