Ovarian Carcinoma Histotype: Strengths and Limitations of Integrating Morphology With Immunohistochemical Predictions
Bibliographic record
Abstract
Ovarian carcinoma histotypes are critical for research and patient management and currently assigned by a combination of histomorphology +/- ancillary immunohistochemistry (IHC). We aimed to validate the previously described IHC algorithm (Calculator of Ovarian carcinoma Subtype/histotype Probability version 3, COSPv3) in an independent population-based cohort, and to identify problem areas for IHC predictions. Histotype was abstracted from cancer registries for eligible ovarian carcinoma cases diagnosed from 2002 to 2011 in Alberta and British Columbia, Canada. Slides were reviewed according to World Health Organization 2014 criteria, tissue microarrays were stained with and scored for the 8 COSPv3 IHC markers, and COSPv3 histotype predictions were calculated. Discordant cases for review and COSPv3 prediction were arbitrated by integrating morphology with IHC results. The integrated histotype (N=880) was then used to identify areas of inferior COSPv3 performance. Review histotype and integrated histotype demonstrated 93% agreement suggesting that IHC information revises expert review in up to 7% of cases. There was also 93% agreement between COSPv3 prediction and integrated histotype. COSPv3 errors predominated in 4 areas: endometrioid carcinoma (EC) versus clear cell (N=23), EC versus low-grade serous (N=15), EC versus high-grade serous (N=11), and high-grade versus low-grade serous (N=6). Most problems were related to Napsin A-negative clear cell, WT1-positive EC, and p53 IHC wild-type high-grade serous carcinomas. Although 93% of COSPv3 prediction accuracy was validated, some histotyping required integration of morphology with ancillary test results. Awareness of these limitations will avoid overreliance on IHC and misclassification of histotypes for research and clinical management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".