Predicting the risk of developing oropharyngeal cancer for Canadians: current evidence and models
Bibliographic record
Abstract
Background: Every year, more than half a million people are diagnosed with Head and Neck Cancers (HNCs). Among different cancers, HNC has a high mortality and morbidity rate. While the etiology of HNC has been known for many years, there has been a rise in the incidence of a subset of these cancers, mainly oropharyngeal cancer, in high income countries including Canada over the past decades. A considerable part of this rise has been attributed to the human papilloma virus (HPV). Therefore, preventive interventions such as vaccination against HPV infection are expected to reduce the number of new oropharyngeal cancer cases. To have efficient prevention, the interventions need to be targeted at high-risk individuals. Risk prediction models can improve the efficiency of these preventive programs by estimating the individualized risk of developing HNC and identifying the high-risk population. Different risk prediction models have been developed worldwide; however, little is known about these models and their applicability in the Canadian context.Objectives: This thesis aims to: 1) review the literature on the HNC risk prediction models and 2) validate a risk prediction model on a sample of the Canadian population.Methods: First, we reviewed the published articles on HNC risk prediction modeling. We included the full-text of peer-reviewed publications that reported at least one model for predicting the risk of developing HNC. We only considered the models that can be used in the primary clinical settings, thus, excluded the ones with genetic markers. This review identified a model that was potentially applicable to the Canadian context. The model was developed to predict the one-year risk of developing oropharyngeal cancer in the US population. In the second step of this thesis project, we validated the predictions of this model on the dataset derived from the Canadian site of the HeNCe Life study, a case-control investigation on the etiology of HNC through a life-course framework in Canada. Based on the model’s development study, we derived a dataset from HeNCe Life comprising 214 cases of oropharyngeal cancers and 433 controls, frequency matched to the cases by sex and 5-year age categories. We replicated the model and tested its predictions on the derived dataset. We evaluated the model’s overall prediction performance by measuring Somers’ D, Brier scores, and R2. The discrimination ability was tested using C-Statistics and discrimination indices. The model’s calibration was assessed by evaluating the calibration slope and intercept values. Results: The first step of this thesis identified nine peer-reviewed HNC risk prediction modeling studies that overall reported 16 models. Six of these studies were conducted in Asia, and only three were published from Western countries, but none from used Canadian data. Most of the models were developed by multivariable logistic regression analysis. All included studies had a high risk of bias, and two of them had high concerns about applicability of the models. Although we did not identify any article reporting a development or validation of a model for the Canadian population, the review found an oropharyngeal cancer risk prediction model, developed in a sample of the US population, that is reproducible and potentially applicable in the Canadian context. Its predictors comprised age, sex, race, pack-years of smoking, previous year’s alcohol consumption, number of lifetime sexual partners, oral HPV infection status, and two-way interaction between sex, pack-years of smoking, and oral HPV infection status. In summary, although the model showed a moderately high level of discrimination, it had poor calibration performance.Conclusion: Limited numbers of HNC risk prediction modeling studies provide sufficient information to judge the models’ quality and applicability. However, the review identified one model that may still be used in the Canadian context
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.044 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.007 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.001 |
| Open science | 0.005 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".