Accuracy and False-Positive Rate of the Cytologic Diagnosis of Follicular Cervicitis: Observations From the College of American Pathologists Pap Educational Program
Bibliographic record
Abstract
CONTEXT: Follicular cervicitis is usually easily identifiable on Papanicolaou (Pap) tests; however, historically, follicular cervicitis is reported to lead to false-positive diagnoses of epithelial cell abnormalities. OBJECTIVES: To assess participant responses in the College of American Pathologists (CAP) Pap educational program (CAP-PAP) to determine the accuracy and false-positive rate of follicular cervicitis cases. Design.-We performed a retrospective review of 4914 participant responses for gynecologic cytology challenges with the reference diagnosis of follicular cervicitis during 11 years (2000-2010) from CAP-PAP. Reference diagnosis category, false-positive rates by participant type (laboratory, cytotechnologist, pathologist), and preparation type (conventional smears, ThinPrep) were analyzed. RESULTS: Of the total 4914 general category responses, 4368 (88.9%) were benign while 546 (11.1%) responses were epithelial cell abnormalities (false positives). Of benign responses, only 2026 (46.4%) were an exact match to follicular cervicitis. Adenocarcinoma and high-grade squamous intraepithelial lesion were the most common diagnoses chosen as a false-positive interpretation (42.3% and 20.1%, respectively). Participant type was significantly associated with false-positive interpretations (laboratory: 19.2%; cytotechnologist: 11.1%; pathologist: 7.9%; P < .001). ThinPrep was also significantly associated with false-positive results as compared to conventional smears (12.2% versus 3.6%; P < .001). CONCLUSIONS: In an interlaboratory comparison educational program, follicular cervicitis is difficult to interpret accurately and represents an important cause of false-positive responses. Follicular cervicitis may mimic adenocarcinoma or high-grade squamous intraepithelial lesion, particularly in liquid-based preparations. The diagnostic difficulty most likely arises from the lymphocytes being less conspicuous in the background as well as their tendency to aggregate in ThinPrep as compared to conventional smears.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.037 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".