190: Pediatric Elbow Fractures: Diagnostic Accuracy of the Combination of Point Tenderness with the Elbow Extension Test
Bibliographic record
Abstract
Elbow injuries are common in children and adolescents accounting for 3% of total emergency room visits. The elbow extension test alone has not been shown to reliably exclude the presence of fractures in pediatrics, thereby most children are referred for a diagnostic elbow x-ray. Determine whether the combination of point tenderness and the elbow extension test improves the diagnostic accuracy for pediatric elbow injuries requiring immobilization. Prospective observational study of patients less than 18 years of age with acute elbow injuries presenting to a tertiary care pediatric emergency department. Patients were managed according to the discretion of the treating physician with imaging and immobilization. The index test was defined as positive if elbow tenderness was present at one of five locations or had incomplete extension of the injured elbow. The assessment was performed by attending physicians or senior residents (≥PGY-3) and inter-observer reliability was assessed on ∼10%. The radiologist's report of the radiographs or, for those without radiographs, a diagnosis of a missed fracture after a structured follow-up phone call one week post-injury was used as the reference standard for the diagnosis of a fracture and/or elbow effusion. A total of 2331 patients were screened with elbow injuries; 1380 excluded (66% pulled elbow, 22% referred in, 7% obvious deformity, 2% uncooperative with exam, 2% polytrauma and 1% other). Amongst the 951 eligible patients, a research assistant recruited a convenience sample of 43%. The mean age was 8.8±4.1 years and 52% male. Of the 334 recruited patients, 97.9% had radiographs and 186 (57%) were diagnosed with an elbow injury requiring immobilization (75.3% had a fracture, 24.7% an isolated effusion). A positive index test was present in 302 (90.4%) and had a sensitivity and specificity of 95.6% (95% CI 91.0% to 98.1%) and 16.9% (95% CI 11.1% to 24.7%), respectively. In comparison, elbow extension alone had a sensitivity and specificity of 82.1% (95% CI 75.1% to 87.5%) and 62.3% (95% CI 53.3% to 70.5%), respectively. Inter-observer reliability assessed on ∼10% was excellent (kappa of 1 for index test and 0.86±0.09 for elbow extension test). Diagnostic accuracy for elbow fractures as the reference standard (ie, not isolated effusions) revealed a sensitivity of 99.2% (95% CI 97.7% to 100.7%) and specificity 16.3% (95% CI 10.5% to 22.0%). Subgroup analysis revealed similar trends for children younger than three years of age. The addition of point tenderness to the elbow extension test improves the sensitivity for the detection of elbow injuries requiring imaging and immobilization, but at a significant decrement of specificity. Future studies are needed to determine how to improve the detection of injuries requiring immobilization from those who can forego any further diagnostic tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".