Missing the Mark: Inaccuracy of administrative data in identification of hospitalized patients with pneumonia and results of a systematic clinical reclassification process on readmission rates
Bibliographic record
Abstract
Objective: Pneumonia readmissions carry financial ramifications under the Hospital Readmissions Reduction Program (HRRP). As readmission determination utilizes administrative data, healthcare systems should evaluate accuracy of pneumonia diagnoses. We sought to develop a systemic process for pneumonia classification review and determine potential effects on pneumonia readmissions in a tertiary academic medical center in the United States.Methods: We performed independent reviews of all pneumonia discharges within 48 hours of discharge over a one-year period. We reclassified all pneumonia discharges into four categories based on the Centers for Disease Control and Prevention reference standard. Secondary review of discordant classifications was performed by discharging providers to determine final diagnosis. The primary outcome was readmission rate within 30 days by pneumonia clinical classification category.Results: Two hundred seventy-eight discharges were reviewed, with overall readmission rate of 18.0%. Independent review confirmed 191 cases (68.7%) as definite or probable pneumonia, while 87 cases (31.3%) were classified as either probably not or not pneumonia. Readmission rates differed significantly between cases reviewed as pneumonia vs. those reviewed as unlikely to be pneumonia (14.1% vs. 26.4%, p < .02). Discharging attending physicians agreed with independent reviewers in 58/87 cases (66.6%), attenuating readmission differences (rate 16.8% for those finalized as pneumonia vs. 22.4% for another diagnosis, p = .32). Pneumonia readmissions were reduced by 1.2% using the classification standard.Conclusions: Complex conditions such as pneumonia may be inaccurately diagnosed in many patients, potentially affecting penalties associated with readmission rates. Therefore, it is imperative that healthcare systems adopt systematic review processes to standardize diagnoses and improve comparative administrative data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".