Reliability and Agreement of Ultrasonographic Measures of the Ovarian Stroma: Impact of Methodology
Bibliographic record
Abstract
OBJECTIVES: Increased ovarian stromal area (SA), stromal-to-ovarian area ratio (S/A), and echogenicity (SEcho) on ultrasonography have been proposed as diagnostic markers for polycystic ovary syndrome. Although several methods to evaluate the stroma exist, their reproducibility has not been defined which limits clinical utility. This study aimed to determine the interrater reliability and agreement of methods to evaluate SA, S/A, and SEcho. METHODS: Five raters tested 3 methods to obtain SA and S/A, and one to obtain SEcho on 30 ovarian cineloops under two imaging conditions, simulating real-time (free-choice) or offline (fixed-frame) imaging. For SA, Method 1 subtracted follicular area from the ovarian area, Method 2 involved outlining the periphery of the stroma, and Method 3 represented a hybrid approach in which central follicles were subtracted from the outlined stroma. SEcho was scored on a subjective 3-tiered scale. Intraclass correlation coefficients (ICCs) and the coefficient of variation were determined for SA and S/A, and Fleiss' kappa agreement statistics (κ) were determined for SEcho. RESULTS: Interrater reliability of SA was superior using Method 1 (ICC = 0.558 and ICC = 0.705) versus Method 2 (ICC = 0.522 and ICC = 0.230) or Method 3 (ICC = 0.429 and ICC = 0.305) under free-choice and fixed-frame imaging conditions, respectively. Interrater reliability of S/A was also moderate to poor across methods. SEcho was also not reliably assessed across raters (κ = <0.500). CONCLUSIONS: Ultrasonographic assessments of the ovarian stroma were associated with moderate to poor reproducibility. Indirect estimates of the ovarian stroma (Method 1) could be optimized to yield a reproducible approach, clarifying the clinical relevance of the stroma.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".