Atypical ductal hyperplasia: interobserver and intraobserver variability, Modern Pathology: An Official Journal of the United States and Canadian Academy of Pathology
Bibliographic record
Abstract
Interobserver reproducibility in the diagnosis of benign intraductal proliferative lesions has been poor. The aims of the study were to investigate the inter- and intraobserver variability and the impact of the addition of an immunostain for high- and low-molecular weight keratins on the variability. Nine pathologists reviewed 81 cases of breast proliferative lesions in three stages and assigned each of the lesions to one of the following three diagnoses: usual ductal hyperplasia, atypical ductal hyperplasia and ductal carcinoma in situ. Hematoxylin and eosin slides and corresponding slides stained with ADH-5 cocktail (cytokeratins (CK) 5, 14. 7, 18 and p63) by immunohistochemistry were evaluated. Concordance was evaluated at each stage of the study. The interobserver agreement among the nine pathologists for diagnosing the 81 proliferative breast lesions was fair (j-value 0.34). The intraobserver j-value ranged from 0.56 to 0.88 (moderate to strong). Complete agreement among nine pathologists was achieved in only nine (11%) cases, at least eight agreed in 20 (25%) cases and seven or more agreed in 38 (47%) cases. Following immunohistochemical stain, a significant improvement in the interobserver concordance (overall j-value 0.50) was observed (P 0.015). There was a significant reduction in the total number of atypical ductal hyperplasia diagnosis made by nine pathologists after the use of ADH-5 immunostain. Atypical ductal hyperplasia still remains a diagnostic dilemma with wide variation in both inter- and intraobserver reproducibility among pathologists. The addition of an immunohis-
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.063 | 0.066 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".