Online Training and Self-assessment in the Histopathologic Classification of Endocervical Adenocarcinoma and Diagnosis of Pattern of Invasion: Evaluation of Participant Performance
Bibliographic record
Abstract
Histopathologic classification of endocervical adenocarcinomas (EAC) has recently changed, with the new system based on human papillomavirus (HPV)-related morphologic features being incorporated into the 5th edition of the WHO Blue Book (Classification of Tumours of the Female Genital Tract). There has also been the introduction of a pattern-based classification system to assess invasion in HPV-associated (HPVA) endocervical adenocarcinomas that stratifies tumors into 3 groups with different prognoses. To facilitate the introduction of these changes into routine clinical practice, websites with training sets and test sets of scanned whole slide images were designed to improve diagnostic performance in histotype classification of endocervical adenocarcinoma based on the International Endocervical Adenocarcinoma Criteria and Classification (IECC) and assessment of Silva pattern of invasion in HPVA endocervical adenocarcinomas. We report on the diagnostic results of those who have participated thus far in these educational websites. Our goal was to identify areas where diagnostic performance was suboptimal and future educational efforts could be directed. There was very good ability to distinguish HPVA from HPV-independent adenocarcinomas within the WHO/IECC classification, with some challenges in the diagnosis of HPV-independent subtypes, especially mesonephric carcinoma. Diagnosis of HPVA subtypes was not consistent. For the Silva classification, the main challenge was related to distinction between pattern A and pattern B, with a tendency for participants to overdiagnose pattern B invasion. These observations can serve as the basis for more targeted efforts to improve diagnostic performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".