Reliability of 3 Radiologic Classifications for the Severity of the Developmental Dysplasia of the Hip in Children Older Than 4 Years
Bibliographic record
Abstract
BACKGROUND: Tonnis, International Hip Dysplasia Institute (IHDI), and lateral metaphyseal height (LMH) are commonly used classifications for grading the severity of the developmental dysplasia of the hip. The reliability of these classifications is not widely studied in older children. The aim of the study was to evaluate the reliability of these 3 radiologic classifications in children older than 4 years and compared with children younger than 4 years and evaluate the cases with varied inter-rater reliability. METHODS: A purposeful sample of 40 children with untreated developmental dysplasia of the hip with ages between 6 months to 8 years was studied for the assessment of the severity grading according to all 3 classifications. Six pediatric orthopaedic surgeons classified all hips for all 3 categorical classifications as per the original description. Inter-rater and intrarater reliability was calculated according to the intraclass correlation coefficient. The cases with different ratings were assessed in detail to evaluate the reasons for the varied rating. RESULTS: The interobserver and intraobserver reliability of all 3 classifications were excellent [intraclass correlation coefficient (ICC): 0.935, 0.820, and 0.935 for IHDI, Tonnis, and LMH classification, respectively]. The excellent reliability was also observed in younger and older children. Interobserver reliability of only dysplastic hips (52 hips) was good for Tonnis (ICC: 0.741) and excellent for IHDI (ICC: 0.911) and LMH classification (ICC-0.9). The main reason for the varied rating was because of the varied perception of the superolateral margin of the acetabulum in few hips. CONCLUSION: The inter-rater and intrarater reliability of all 3 classifications (IHDI, Tonnis, and LMH) is excellent. All classifications can be used till the age of 8 years. The difficulty in selecting the superolateral margin of the acetabulum is a major cause of inter-rater variability. LEVEL OF STUDY: Level III.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".