Trends in Results of HBIPS National Performance Measures and Association With Year of Adoption
Bibliographic record
Abstract
OBJECTIVE: Multiple studies demonstrate a consistent pattern of improvement on quality measures among health care organizations after they begin collecting and reporting data. This study compared results on psychiatric performance measures among cohorts of hospitals with different characteristics that elected to begin reporting on the measures at various points in time. METHODS: Quarterly reporting of Hospital-Based Inpatient Psychiatric Services (HBIPS) measures to the Joint Commission was used to examine trends in performance among four hospital cohorts that began reporting in 2009 (N=243), 2011 (N=139), 2014 (N=137), or 2015 (N=372). The HBIPS measures address admission screening, restraint and seclusion use, justification of use of multiple antipsychotic medications, and discharge planning. Comparisons were based upon initial quarters of data reported and change rates. RESULTS: After adjustment for covariates, the analyses showed that all cohorts significantly improved across quarters for admission screening, justification of multiple antipsychotic medications, and discharge planning. Restraint hours significantly dropped over the initial reporting periods, but only for the 2009 and 2015 cohorts. Seclusion hours significantly dropped over the six reporting periods for all cohorts except 2011. CONCLUSIONS: Several differences were observed across cohorts in the rate of change between baseline and final measurement for various measures. In nearly every case, however, hospitals that began reporting measurement data earlier performed better than subsequent cohorts during the later cohorts' first quarter of reporting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".