Gleason grade 4 prostate adenocarcinoma patterns: an interobserver agreement study among genitourinary pathologists
Bibliographic record
Abstract
AIMS: To assess the interobserver reproducibility of individual Gleason grade 4 growth patterns. METHODS AND RESULTS: Twenty-three genitourinary pathologists participated in the evaluation of 60 selected high-magnification photographs. The selection included 10 cases of Gleason grade 3, 40 of Gleason grade 4 (10 per growth pattern), and 10 of Gleason grade 5. Participants were asked to select a single predominant Gleason grade per case (3, 4, or 5), and to indicate the predominant Gleason grade 4 growth pattern, if present. 'Consensus' was defined as at least 80% agreement, and 'favoured' as 60-80% agreement. Consensus on Gleason grading was reached in 47 of 60 (78%) cases, 35 of which were assigned to grade 4. In the 13 non-consensus cases, ill-formed (6/13, 46%) and fused (7/13, 54%) patterns were involved in the disagreement. Among the 20 cases where at least one pathologist assigned the ill-formed growth pattern, none (0%, 0/20) reached consensus. Consensus for fused, cribriform and glomeruloid glands was reached in 2%, 23% and 38% of cases, respectively. In nine of 35 (26%) consensus Gleason grade 4 cases, participants disagreed on the growth pattern. Six of these were characterized by large epithelial proliferations with delicate intervening fibrovascular cores, which were alternatively given the designation fused or cribriform growth pattern ('complex fused'). CONCLUSIONS: Consensus on Gleason grade 4 growth pattern was predominantly reached on cribriform and glomeruloid patterns, but rarely on ill-formed and fused glands. The complex fused glands seem to constitute a borderline pattern of unknown prognostic significance on which a consensus could not be reached.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.059 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".