Concordance of performance metrics among US trauma centers caring for injured children
Bibliographic record
Abstract
BACKGROUND: Several indicators of quality pediatric trauma care have been proposed including low in-hospital mortality, nonoperative management of blunt splenic injury, use of intracranial pressure monitors after severe traumatic brain injury, and craniotomy for children with severe subdural or epidural hematomas. It is not known if center-level performance is consistent in each of these metrics. We evaluated whether center performance in one area of quality predicted similar performance in other areas of quality. METHODS: We reviewed patients 18 years or younger who were hospitalized with an injury Abbreviated Injury Scale (AIS) score of 2 or greater from 2010 to 2011 at trauma centers (n = 150) participating in the Trauma Quality Improvement Program. Random-intercept multilevel modeling was used to generate center-specific adjusted odds ratios for each quality indicator. We evaluated correlations between center-specific adjusted odds ratios of each quality indicator and mortality using Pearson correlation coefficients. Weighted κ statistics were used to test multiple pairwise agreements between indicators and the overall agreement across all four indicators. RESULTS: Among 84,880 children identified for analysis, 3,603 had blunt splenic injury, 3,503 had severe traumatic brain injury, and 1,286 had an epidural or subdural hematoma. A negative correlation between center-specific odds of mortality and craniotomy was present (Pearson correlation coefficient, -0.18; p = 0.03). There were no significant correlations between other indicators. Although κ statistics showed slight agreement for the pairwise comparison of odds of mortality and craniotomy (0.17, 0.02-0.32), there was no agreement for all other pairwise comparisons or the overall comparison of all four indicators (-0.01, -0.07 to 0.06). CONCLUSION: Our findings demonstrate a lack of concordance in center-level performance across the four pediatric trauma quality indicators we evaluated. These findings should be considered by pediatric trauma quality improvement initiatives to allow for comprehensive measurement of hospital quality as opposed to benchmarking using a single indicator.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.071 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".