Comparison of Performance of Conventional and ThinPrep Gynecologic Preparations in the College of American Pathologists Gynecologic Cytology Program
Bibliographic record
Abstract
Abstract Context. —Results of clinical trials suggest that interpretation of liquid-based cytology preparations is more accurate and is associated with less screening error than interpretation of conventional preparations. Objective. —In this study, the performance of participants in interpreting ThinPrep (TP) preparations was compared with participants' performance on conventional Papanicolaou tests in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology (PAP). Design. —The results of the PAP from the year 2002 were reviewed, and the discordancies to series and exact-match error rates for the 2 cytologic methods were compared. Results. —For this study, a total of 89 815 interpretations from conventional smears and 20 886 interpretations from TP samples were analyzed. Overall, interpretations of TP preparations had both significantly fewer false-positive (1.6%) and false-negative (1.3%) rates than those of conventional smears ( P = .001 and P = .02, respectively) for validated or validated-equivalent slides, as assessed by concordance with the correct diagnostic series. In this assessment of concordance to series, interpretations of educational TP and conventional preparations were similar, except for high-grade squamous intraepithelial lesion, in which the performance was significantly worse for educational TP preparations (false-negative rate of 8.1% vs 4.1% for conventional smears, P < .001). When interpretations were matched to the exact diagnosis, validated-equivalent TP preparations were generally more accurate for diagnoses in the 100 series and 200 series than were conventional smears. Notably, for the reference diagnosis of squamous cell carcinoma, the exact-match error rate on validated equivalent TP slides was significantly greater than that of conventional slides (44.5% vs 23.1%, P < .001). Interpretations of educational TP preparations also had a significantly higher error rate in matching to the exact reference diagnosis for squamous cell carcinoma (33.7% vs 22.8%, P = .007). Conclusions. —Overall, TP preparations in this program were associated with significantly lower error rates than conventional smears for both validated and educational cases. However, unlike the negative for intraepithelial lesion and malignancy, not otherwise specified, low-grade squamous intraepithelial lesion, and adenocarcinoma cytodiagnostic challenges, participants' responses indicated some difficulty in recognizing high-grade squamous intraepithelial lesion and squamous cell carcinoma.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.006 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".