Challenges of false positive and negative results in cervical cancer screening
Bibliographic record
Abstract
ABSTRACT Objective To quantify the impact and accuracy of different screening approaches for cervical cancer, including liquid based cytology (LBC), molecular testing for human papillomavirus (HPV) infection, and their combinations via parallel co-testing and sequential triage. The secondary goal was to predict the effect of differing coverage rates of HPV vaccination on the performance of screening tests and in the interpretation of their results. Design Modelling study. Main outcomes measured Different screening modalities were compared in terms of number of cases of Cervical intra-epithelial neoplasia (CIN) grade 2 and 3 detected and missed, as well as the number of false positives leading to excess colposcopy, and number of tests required to achieve a given level of accuracy. The positive predictive value (PPV) and negative predictive value (NPV) of different modalities were simulated under varying levels of HPV vaccination. Results The model predicted that in a typical population, primary LBC screening misses 4.9 (95% Confidence Interval (CI) 3.5-CIN 2 / 3 cases per 1000 women, and results in 95 (95% CI: 93-97%) false positives leading to excess colposcopy. For primary HPV testing, 2.0 (95% CI: 1.9-2.1) cases were missed per 1000 women, with 99 (95% CI: 98-101) excess colposcopies undertaken. Co-testing markedly reduced missed cases to 0.5 (95% CI: 0.3-0.7) per 1000 women, but at the cost of dramatically increasing excess colposcopy referral to 184 per 1000 women (95% CI: 182-188). Conversely, triage testing with reflex screening substantially reduced excess colposcopy to 9.6 cases per 1000 women (95% CI: 9.3 - 10) but at the cost of missing more cases (6.4 per 1000 women, 95% CI: 5.1 - 8.0). Over a life-time of screening, women who always attend annual and 3-year co-testing were predicted to have a virtually 100% chance of falsely detecting a CIN 2 / 3 case, while 5 year co-testing has a 93.8% chance of a false positive over screening life-time. For annual, 3 year, and 5 year triage testing (either LBC with HPV reflex or vice versa), lifetime risk of a false positive is 35.1%, 13.4%, and 8.3% respectively. HPV vaccination rates adversely impact the PPV, while increasing the NPV of various screening modalities. Results of this work indicate that as HPV vaccination rates increase, HPV based screening approaches result in fewer unnecessary colposcopies than LBC approaches. Conclusion The clinical relevance of cervical cancer screening is crucially dependent upon the prevalence of cervical dysplasia and/or HPV infection or vaccination in a given population, as well as the sensitivity and specificity of various modalities. Although screening is life-saving, false negatives and positives will occur, and over-testing may cause significant harm, including potential over-treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.136 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.005 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".