A Comparison of 2 Abbreviated Methods for Assessing Adolescent Bone Age: The Shorthand Bone Age Method and the SickKids/Columbia Method
Bibliographic record
Abstract
BACKGROUND: Radiographic assessment of bone age is critically important to decision-making on the type and timing of operative interventions in pediatric orthopaedics. The current widely accepted method for determining bone age is time and resource-intensive. This study sought to assess the reliability and accuracy of 2 abbreviated methods, the Shorthand Bone Age (SBA) and the SickKids/Columbia (SKC) methods, to the widely accepted Greulich and Pyle (GP) method. METHODS: Standard posteroanterior radiographs of the left hand of 125 adolescent males and 125 adolescent females were compiled, with bone ages determined by the GP method ranging from 9 to 16 years for males and 8 to 14 years for females. Blinded to the chronologic age and GP bone age of each child, the bone age for each radiograph was determined using the SBA and SKC methods by an orthopaedic surgery resident, 2 pediatric orthopaedic surgeons, and a musculoskeletal radiologist. Measurements were then repeated 2 weeks later after rerandomization of the radiographs. Intrarater and interrater reliability for the 2 abbreviated methods as well as the agreement between all 3 methods were calculated using weighted κ values. Mean absolute differences between methods were also calculated. RESULTS: Both bone age methods demonstrated substantial to almost perfect intrarater reliability, with a weighted κ ranging from 0.79 to 0.93 for the SBA method and from 0.82 to 0.96 for the SKC method. Interrater reliability was moderate to substantial (weighted κ: 0.55 to 0.84) for the SBA method and substantial to almost perfect (weighted κ: 0.67 to 0.92) for the SKC method. Agreement between the 3 methods was substantial for all raters and all comparisons. The mean absolute difference, been GP-derived and SBA-derived bone age, was 7.6±7.8 months, as compared with 8.8±7.4 months between GP-derived and SKC-derived bone ages. CONCLUSIONS: The SBA and SKC methods have comparable reliability, and both correlate well to the widely accepted GP methods and to each other. However, they have relatively large absolute differences when compared with the GP method. These methods offer simple, efficient, and affordable estimates for bone age determination, but at best provide an estimate to be used in the appropriate setting. LEVEL OF EVIDENCE: Diagnostic study-level III.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".