Papanicolaou Tests With Mixed High-Grade and Low‐Grade Squamous Intraepithelial Lesion Features: Distinct Performance in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytopathology
Bibliographic record
Abstract
Abstract Context. —Previous studies have shown that in gynecologic cytology, cases of low-grade squamous intraepithelial lesion (LSIL) and high-grade squamous intraepithelial lesion (HSIL) perform differently on interpretive review. The performance of cases with mixed LSIL and HSIL features is unknown. Objective. —To compare the performance of gynecologic cytology cases of “pure” LSIL and HSIL with cases showing mixed LSIL and HSIL features. Design. —We compiled performance data from the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytopathology from the years 2003 and 2004, and compared the performance of slides showing relatively pure LSIL and HSIL (≤10% misclassification as HSIL and LSIL, respectively) with slides showing mixed LSIL or HSIL features (cases misclassified as LSIL or HSIL >10% of the time). Results. —Interpretations from a total of 4508 cases (2452 HSIL and 2056 LSIL) were analyzed. Overall, the sensitivity of participants on slides with a reference diagnosis of HSIL was 97.3%, and of LSIL was 95.9%. Performance trends for pure versus mixed cases varied by slide type and reference diagnosis. For conventional slides, participant sensitivity on pure HSIL cases was greatest (98.0%) and on pure LSIL cases was least (95.2%), while participant performance on cases with mixed features was intermediate (97.0% for mixed HSIL and 96.7% for mixed LSIL). In contrast, participant performance on ThinPrep slides showed the greatest sensitivity for mixed LSIL slides (97.9%), while performance on mixed HSIL slides showed the lowest sensitivity (95.7%); slides with pure features had intermediate sensitivity levels (96.3% for both HSIL and LSIL). Further evaluation demonstrated that conventional pure HSIL slides performed significantly better than mixed HSIL slides ( P = .006), whereas mixed LSIL slides performed better than pure LSIL slides ( P = .01). For ThinPrep slides, pure HSIL cases performed similarly to mixed HSIL cases ( P = .43), while mixed LSIL cases performed better than pure LSIL cases ( P = .04). Conclusion. —Slides with mixed LSIL and HSIL features have measurably distinct performance characteristics in comparison to slides with pure LSIL or HSIL features. Participant performance on conventional mixed cases is distinctly different from performance on ThinPrep mixed cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.043 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".