A Systematic Review and Meta-analysis on the Reproducibility of Ultrasound-based Metrics for Assessing Developmental Dysplasia of the Hip
Bibliographic record
Abstract
BACKGROUND: The purposes of this study were to (1) perform a systematic review of articles that reported agreement or reproducibility in repeated diagnosis of developmental dysplasia of the hip (DDH) using ultrasound imaging, (2) estimate the reproducibility in the available dysplasia metrics, and (3) compare reproducibility of the available dysplasia metrics. METHODS: A systematic review of the Medline and Embase databases was performed by using a search strategy formulated from our research question: "For infants at risk of DDH, are US imaging-based diagnoses reproducible?" Two reviewers independently identified articles for inclusion in the systematic review, and then assessed the quality of the included studies using the Guidelines for Reporting Reliability and Agreement Studies guideline. Variability and agreement-related statistics in the included studies were extracted and included in a meta-analysis for summarizing the available statistics. The reproducibility of the available dysplasia metrics was compared, with a Bonferroni correction made to adjust for multiple comparisons. RESULTS: Twenty eight studies were included in the systematic review. Overall, the quality of the included studies was moderate (average, 10.7/15; range, 6 to 12). Graf's alpha angle had the lowest interexamination variability of the metrics assessed, followed by Graf's beta angle (the variability of the alpha angle was 10% lower than the variability of the beta angle, P<0.05). However, despite Graf's angles having lower variability compared with other dysplasia metrics, their actual variability was still problematically high. This finding was supported by the low intraclass correlation and Kappa coefficient values reported in the included studies. There was also evidence to suggest that the reproducibility in DDH diagnosis has potentially worsened over time. CONCLUSIONS: Overall, we found high variability and low agreement in all reported dysplasia metrics. Furthermore, in the last 3 decades, the repeatability of dysplasia metrics has not markedly improved and may even have declined, indicating a genuine need for improving repeatability and reliability of ultrasound-based DDH diagnosis. LEVEL OF EVIDENCE: Level III-systematic review of level III studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.058 | 0.156 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.022 | 0.041 |
| Bibliometrics | 0.013 | 0.012 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".