Measuring the Significance of Participant Evaluation of Acceptability of Cases in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology
Bibliographic record
Abstract
CONTEXT: The quality of gynecologic cytology slides within educational and proficiency testing programs may deteriorate during use. Participant evaluation of the acceptability of these slides subsequent to possible deterioration is not known. OBJECTIVE: To assess participants' evaluation of the acceptability of slides circulating within an educational gynecologic cytology glass slide program. DESIGN: The College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology is a peer comparison and educational program that evaluates the ability of the participants to correctly classify gynecologic cytology preparations. The program uses both expert review and field validation to select slides for the graded portion of the program. Participants were asked to assess the acceptability of slides within the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology, and their responses were assessed with respect to type of slide preparation, validation status, and reference diagnosis. In addition, we compared the cytodiagnostic discordancy rates of slides that were deemed acceptable by participants with those that were deemed unacceptable. SETTING: Participant assessments were derived from pathologists and cytotechnologists from cytology laboratories of all types. RESULTS: A total of 17,210 slide interpretations were reviewed, and 2.91% of the cases were labeled unacceptable by participants. For all slides, the percentage of cases called unacceptable varied from 1.65% for cases with a reference interpretation of herpes to 45% for cases with a reference interpretation of unsatisfactory. The percentage of slides deemed unacceptable was higher for validated slides than for educational slides (3.27% vs 2.55%, P = .006). The discordancy rate (to reference diagnosis series) for cases deemed unacceptable was significantly higher than the discordancy rate for cases deemed acceptable for both validated (10.39% vs 1.76%) and educational slides (21.72% vs 3.53%, P < .001). CONCLUSION: Greater than 97% of all slides in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology were judged acceptable by participants. Despite expert review and field validation, a small percentage of slides (almost 3%) in this program were deemed unacceptable by participants. These results support the use of participant evaluation of cases to continually improve the quality of cases in this program.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.106 | 0.232 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".