Reliability of Indices Measured on Infant Hip MRI at Time of Spica Cast Application for Dysplasia
Bibliographic record
Abstract
PURPOSE: Infants with persistent developmental dysplasia of the hip (DDH) after harness treatment may be treated by operative reduction and spica casting, with post-reduction hip alignment assessed by magnetic resonance imaging (MRI), which can demonstrate three-dimensional hip geometry. This may provide valuable information regarding prognosis and adequacy of management, but such scans are difficult to assess due to limited spatial resolution and artefacts caused by patient motion. This may account for the limited success in correlating MRI findings to clinical outcomes to date. As a first step to improving these results we tested whether MRI indices of hip deformity and quality of femoral head reduction could be reliably measured. PROCEDURES: We retrospectively studied children with DDH, post-spica-cast MRI, and radiographic follow-up. We measured MRI indices adapted from other reports using computed tomography (CT) and MRI, and added new indices. Inter-observer reliability and inter-index correlations were evaluated and indices adapted during the process. FINDINGS: We observed 55 dysplastic hips in 41 infants. Despite difficulties inherent to infant hip MRI, several indices were measured with substantial agreement (kappa up to 0.88, intra-class correlation ICC up to 0.91), with highly significant (p<0.01) correlation with each other (up to r = 0.72). Reliable indices included coronal acetabular angle, pulvinar fat-pad thickness, presence of a barrier to reduction, and grades of subluxation and dysplasia. CONCLUSION: Indices on MRI scan post hip spica cast placement can be measured reliably, assessing acetabular geometry, degree of hip reduction and barriers to reduction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".