Diagnostic accuracy of <scp>HPV16</scp> early antigen serology for <scp>HPV</scp>‐driven oropharyngeal cancer is independent of age and sex
Bibliographic record
Abstract
Abstract A growing proportion of head and neck cancer (HNC), especially oropharyngeal cancer (OPC), is caused by human papillomavirus (HPV). There are several markers for HPV‐driven HNC, one being HPV early antigen serology. We aimed to investigate the diagnostic accuracy of HPV serology and its performance across patient characteristics. Data from the VOYAGER consortium was used, which comprises five studies on HNC from North America and Europe. Diagnostic accuracy, that is, sensitivity, specificity, Cohen's kappa and correctly classified proportions of HPV16 E6 serology, was assessed for OPC and other HNC using p16INK4a immunohistochemistry (p16), HPV in situ hybridization (ISH) and HPV PCR as reference methods. Stratified analyses were performed for variables including age, sex, smoking and alcohol use, to test the robustness of diagnostic accuracy. A risk‐factor analysis based on serology was conducted, comparing HPV‐driven to non‐HPV‐driven OPC. Overall, HPV serology had a sensitivity of 86.8% (95% CI 85.1‐88.3) and specificity of 91.2% (95% CI 88.6‐93.4) for HPV‐driven OPC using p16 as a reference method. In stratified analyses, diagnostic accuracy remained consistent across sex and different age groups. Sensitivity was lower for heavy smokers (77.7%), OPC without lymph node involvement (74.4%) and the ARCAGE study (66.7%), while specificity decreased for cases with <10 pack‐years (72.1%). The risk‐factor model included study, year of diagnosis, age, sex, BMI, alcohol use, pack‐years, TNM‐T and TNM‐N stage. HPV serology is a robust biomarker for HPV‐driven OPC, and its diagnostic accuracy is independent of age and sex. Future research is suggested on the influence of smoking on HPV antibody levels.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".