Diagnosis of Gleason Pattern 5 Prostate Adenocarcinoma on Core Needle Biopsy
Bibliographic record
Abstract
Accurate recognition of Gleason pattern 5 (GP5) prostate adenocarcinoma on needle biopsy is critical as it is associated with disease progression and adverse clinical outcome. Despite important implications of this diagnosis, interobserver variation in the diagnosis of GP5 has not been adequately studied. Digital images of 66 prostate adenocarcinoma cases that potentially contained a GP5 component were distributed to 16 urologic pathologists who were asked to classify whether GP5 was present. Each image was initially classified into 1 of 4 morphologic subpatterns by 2 coauthors (R.B.S. and M.Z.): solid nests (15), comedocarcinoma (8), single cells and/or cords (35), and variant morphology (8). Additional features captured included: size (large: >20 cells, medium: 10 to 20 cells, and small: <10 cells) and distribution of nuclei (uniform vs. nonuniform) for nests pattern; intraluminal coagulative tumor necrosis, karyorrhectic debris, and amorphous material for comedocarcinoma pattern; and quantity (≤5, 6 to 10, and >10) and distribution (clustered vs. intermixed with adjacent well-formed glands) for single cells/cords pattern. Interobserver reproducibility of a diagnosis of GP5 was assessed and the morphologic subpatterns and features were correlated with the consensus diagnosis (defined as 75% agreement). Interobserver reproducibility for overall diagnostic agreement was fair (κ=0.376). Among subpatterns, comedocarcinoma had highest reproducibility (κ=0.499), followed by variant morphology (κ=0.443), single cells/cords (κ=0.369), and nests (κ=0.347). All cases with the following features achieved consensus for GP5: large nests regardless of nuclear distribution; coagulative necrosis with or without karyorrhectic debris; single cells/cords >10 or 6 to 10 in a cluster; and signet ring-like cells in single cells or within nests pattern. A majority of cases with the following features achieved consensus against GP5: medium-size nests; exclusive intraluminal amorphous material; single cells/cords ≤5; and Paneth cell change. Remaining morphologic features did not reach consensus for or against GP5. A majority (86%) of participants would diagnose a small focus of GP5 only when it is present in >1 level. The diagnostic reproducibility of GP5 within certain morphologies was only fair among urologic pathologists. However, the diagnosis of GP5 was more reproducible when certain restrictive morphologic and quantitative criteria were applied. These findings suggest that additional studies are needed to find highly reproducible features of GP5 associated with documented aggressive clinical outcome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".