Comparison of Injury Patient Information From Hospitals With Records in Both the National Trauma Data Bank and the Nationwide Inpatient Sample
Bibliographic record
Abstract
BACKGROUND: Administrative and registry databases are useful for researchers given their availability and size, yet their limitations for specific applications remain undefined. We compared injury records from a large administrative database and the National Trauma Data Bank (NTDB) with the goal of furthering the understanding of their respective limitations. METHODS: The study hospitals had submitted records to both the NTDB and the Nationwide Inpatient Sample (NIS) for patients admitted during 2002. Record inclusion criteria for comparison included nonelective admissions with a primary diagnosis of injury (excluding isolated hip fractures). Numbers of cases and variables common to both databases were compared. RESULTS: Twenty-four hospitals had records both in the NTDB (24,619 records) and in the NIS (25,586 records). We found less missing cost and payer information in the NIS compared with the NTDB (0% and 0.1% vs. 30.5% and 24%, respectively), higher mean number of comorbidities per record in the NIS (0.77 vs. 0.18), and a lower crude case fatality rate in the NIS (3.5% vs. 5.2%). CONCLUSIONS: The main differences between the databases reflected the different motives for data collection and the inclusion or exclusion criteria imposed by trauma registries. These differences require consideration when using either database to investigate injury-related questions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.075 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.016 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".