External Benchmarking of Trauma Center Performance
Bibliographic record
Abstract
OBJECTIVE: The elderly injured have been identified as a population with unique needs compared with nonelderly trauma patients. We sought to determine whether trauma center (TC) performance is consistent across age groups and to assess whether aggregate evaluations of TC performance capture quality of care among the elderly. BACKGROUND: The recently launched Trauma Quality Improvement Program utilizes external benchmarking of TC outcomes to identify centers with above-average performance, with the goal of disseminating best practices. If variation exists in TC performance across age groups, such variation might significantly impact on the success of external benchmarking programs in improving quality of care. METHODS: Study data were derived from the National Trauma Databank (2007), limited to level I and II centers and adults with moderate to severe injuries (injury severity score > 9). Separate logistic regression models were constructed to produce TC risk-adjusted mortality for both the young and the elderly (age > 65 years). Observed-to-expected mortality ratios were used to identify centers with above or below average performance overall, among the young and among the elderly. RESULTS: We identified 87,754 patients across 132 facilities; 25% were elderly. After adjustment for case mix, 9 centers were identified as above-average performers in the elderly population. Only 2 of these centers were also above-average performers among young patients. Overall, concordance for center performance across age strata evidenced poor agreement (κ, 0.23). In addition, aggregate assessment of center performance did not reliably identify high-performing centers for elderly patients. CONCLUSIONS: The use of outcome-based benchmarking harbors significant potential for trauma quality improvement. Evaluations of aggregate TC performance may not adequately reflect the care provided to the elderly injured. Elderly trauma patients may warrant special attention in the context of ongoing quality improvement programs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".