Reproducibility of the Histological Diagnosis of Cervical Dysplasia Among Pathologists From 4 Continents
Bibliographic record
Abstract
The reliable histological diagnosis of cervical squamous intraepithelial lesions (SILs), especially low-grade SIL, is known to be problematic. Poor diagnostic reproducibility can complicate studies addressing its appropriate management. As part of an international study comparing expectant management of histologically proven low-grade SIL with immediate loop electrocautery excisional procedure, this study was carried out to assess interobserver agreement on the histological diagnosis of SILs among a group of 22 pathologists from 5 countries and the intraobserver reliability among a subset of 7 Canadian pathologists. Fifty-six histological slides from colposcopically obtained cervical biopsies were circulated to each of the 22 pathologists. To assess intraobserver reliability, 7 Canadian pathologists assessed 40 of the slides once and 16 of the slides twice. Kappa values were used to measure interobserver agreement with an overall kappa value of 0.61 (95% confidence interval, 0.60-0.62) corresponding to moderate reliability. The weighted kappa values for interobserver agreement ranged from 0.46 to 0.88 (median, 0.79). The intraobserver reliability of 7 Canadian pathologists ranged from substantial to excellent based upon the weighted kappa values ranging from 0.62 to 0.94 (median, 0.72). This degree of reliability is comparable to that found in similar studies. In an individual case, there can be considerable disparity in diagnosis that can result in disparate management strategies. This adds a layer of complexity to any trial that attempts to assess optimal treatment strategies or the natural history of this disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".