Reliability and agreement of manual and automated morphological radiographic hip measurements
Bibliographic record
Abstract
Objective: To determine the reliability and agreement of manual and automated morphological measurements, and agreement in morphological diagnoses. Methods: Thirty pelvic radiographs were randomly selected from the World COACH consortium. Manual and automated measurements of acetabular depth-width ratio (ADR), modified acetabular index (mAI), alpha angle (AA), Wiberg center edge angle (WCEA), lateral center edge angle (LCEA), extrusion index (EI), neck-shaft angle (NSA), and triangular index ratio (TIR) were performed. Bland-Altman plots and intraclass correlation coefficients (ICCs) were used to test reliability. Agreement in diagnosing acetabular dysplasia, pincer and cam morphology by manual and automated measurements was assessed using percentage agreement. Visualizations of all measurements were scored by a radiologist. Results: The Bland-Altman plots showed no to small mean differences between automated and manual measurements for all measurements except for ADR. Intraobserver ICCs of manual measurements ranged from 0.26 (95%-CI 0-0.57) for TIR to 0.95 (95%-CI 0.87-0.98) for LCEA. Interobserver ICCs of manual measurements ranged from 0.43 (95%-CI 0.10-0.68) for AA to 0.95 (95%-CI 0.86-0.98) for LCEA. Intermethod ICCs ranged from 0.46 (95%-CI 0.12-0.70) for AA to 0.89 (95%-CI 0.78-0.94) for LCEA. Radiographic diagnostic agreement ranged from 47% to 100% for the manual observers and 63%-96% for the automated method as assessed by the radiologist. Conclusion: The automated algorithm performed equally well compared to manual measurement by trained observers, attesting to its reliability and efficiency in rapidly computing morphological measurements. This validated method can aid clinical practice and accelerate hip osteoarthritis research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.072 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".