Interobserver Reliability for Identifying Specific Patterns of Placental Injury as Defined by the Amsterdam Classification
Bibliographic record
Abstract
CONTEXT.—: Placental pathology is an essential tool for understanding neonatal illness. The recent Amsterdam international consensus has standardized criteria and terminology, providing harmonized data for research and clinical care. OBJECTIVE.—: To evaluate the interobserver reliability of these criteria between pathologists at different levels of experience using digitally scanned slides from placentas in a birth population including a large proportion of normal deliveries. DESIGN.—: This was a secondary analysis of selected placentas from a large case-control study of placental lesions associated with neonatal encephalopathy. Histologic slides from 80 placentas were digitally scanned and blindly evaluated by 6 pathologists. Interobserver reliability was assessed by positive and negative agreement, Fleiss κ, and interrater correlation coefficients. RESULTS.—: Overall agreement on the diagnosis, grading, and staging of acute chorioamnionitis and villitis of unknown etiology was moderate to good for all observers and good to excellent for a subset of 4 observers. Agreement on the diagnosis and subtyping of fetal vascular malperfusion was poor to fair for all observers and fair to moderate for the subset of 4 pathologists. Agreement on accelerated villous maturation was poor. CONCLUSIONS.—: This study critically evaluates interobserver reliability for lesions defined by the Amsterdam consensus using scanned images with a low frequency of pathologic lesions. Although reliability was good to excellent for inflammatory lesions, lower reliability for vascular lesions emphasizes the need to more explicitly define the specific histologic features and boundaries for these patterns.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.081 | 0.126 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".