Reliability and Agreement of Ultrasonographic Measures of the Ovarian Stroma: Impact of Methodology
Bibliographic record
Abstract
OBJECTIVES: Increased ovarian stromal area (SA), stromal-to-ovarian area ratio (S/A), and echogenicity (SEcho) on ultrasonography have been proposed as diagnostic markers for polycystic ovary syndrome. Although several methods to evaluate the stroma exist, their reproducibility has not been defined which limits clinical utility. This study aimed to determine the interrater reliability and agreement of methods to evaluate SA, S/A, and SEcho. METHODS: Five raters tested 3 methods to obtain SA and S/A, and one to obtain SEcho on 30 ovarian cineloops under two imaging conditions, simulating real-time (free-choice) or offline (fixed-frame) imaging. For SA, Method 1 subtracted follicular area from the ovarian area, Method 2 involved outlining the periphery of the stroma, and Method 3 represented a hybrid approach in which central follicles were subtracted from the outlined stroma. SEcho was scored on a subjective 3-tiered scale. Intraclass correlation coefficients (ICCs) and the coefficient of variation were determined for SA and S/A, and Fleiss' kappa agreement statistics (κ) were determined for SEcho. RESULTS: Interrater reliability of SA was superior using Method 1 (ICC = 0.558 and ICC = 0.705) versus Method 2 (ICC = 0.522 and ICC = 0.230) or Method 3 (ICC = 0.429 and ICC = 0.305) under free-choice and fixed-frame imaging conditions, respectively. Interrater reliability of S/A was also moderate to poor across methods. SEcho was also not reliably assessed across raters (κ = <0.500). CONCLUSIONS: Ultrasonographic assessments of the ovarian stroma were associated with moderate to poor reproducibility. Indirect estimates of the ovarian stroma (Method 1) could be optimized to yield a reproducible approach, clarifying the clinical relevance of the stroma.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.072 | 0.134 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".