Risk stratification models in human papillomavirus-associated oropharyngeal squamous cell carcinoma: The Nova Scotia distribution
Bibliographic record
Abstract
OBJECTIVE: The incidence of oropharyngeal squamous cell carcinoma is increasing with a growing proportion of diagnoses associated with human papillomavirus (p16 + OSCC), which generally confers a favorable prognosis. For these reasons, novel risk stratification models specific to the p16 + OSCC population have recently been proposed to guide future research on treatment de-intensification for appropriate patients. This study aimed to quantify patient risk distribution using multiple published risk models and investigate the hypothesis that the local p16 + OSCC population includes a smaller proportion of low-risk patients due to a high prevalence of concurrent tobacco exposure. METHODS: A retrospective cohort study was performed including patients diagnosed with p16 + OSCC in Nova Scotia between 2011 and 2015. Patient identification was obtained through the CCNS registry and an institutional database. Exclusion criteria included HPV negative status, second primary cases, incomplete data availability, and local recurrence cases. RESULTS: Following exclusion, 117 patients met study criteria. The majority had small primary tumors (70.9% ≤ T2) and advanced nodal status on presentation (60.7% ≥ N2b). Most patients had a positive smoking history (62.4%), with 53.0% of patients having a pack-year history greater than 10 pack-years. In four of the five risk stratification models, the majority of the study population fell into the lowest risk category. The risk stratification distribution of our local population was similar to the populations used to validate the published models, with the largest single category difference being 13.3% (range - 12.3 to + 13.3%). CONCLUSIONS: This is the first study to compare multiple currently published risk stratification models to a local population and address the uncertainty of risk stratification in the Nova Scotian p16 + OSCC population. Despite a high prevalence of concurrent tobacco exposure, the study population was found to be overall low risk, with similar risk compared to model validation populations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".