Applying an implementation science lens to understand physician-level variation in patient length of stay in internal medicine
Bibliographic record
Abstract
Abstract Background & objectives Length of Stay (LoS) is a critical quality metric and focus of improvement efforts in healthcare. Successfully managing LoS depends on understanding the drivers of variation amenable to change. This study aims to (1) characterize physician-level variation in LoS; (2) identify physician actions associated with LoS; and (3) explore the individual-, team-, and hospital-level factors influencing this variation to generate hypotheses for further study. Methods This mixed-methods comparative case study approach examined six General Internal Medicine (GIM) departments in Toronto, Ontario. Physician-level variation in LoS was calculated using a random-intercept negative binomial regression model and sensitivity analysis. Semi-structured interviews and ethnographic observations were conducted and analyzed using the AACTT Framework (Action-Actor-Context-Target-Time), the Consolidated Framework for Implementation Research (CFIR), and the Theoretical Domains Frameworks (TDF). Hospitals with the lowest and highest physician-level variation in LoS were compared. Results Physician-level variation in LoS ranged from 1.7 to 7.0%, which—though modest numerically—represents meaningful differences in physician decision-making not explained by patient complexity, and no significant hospital-level effect was observed. Qualitative analysis from 12 observations and 67 interviews (32 GIM physicians and residents, 35 nurses and other health professionals) identified eight discrete physician actions influencing LoS, along with five individual-level factors and five team- and hospital-level factors. The nature of these factors was different when comparing hospitals with the lowest and highest variation. Organizational culture and perceptions of the patient population shaped physician perceptions of their professional role, while GIM departmental culture, structural characteristics, and communication networks informed physician beliefs about team capabilities and consequences of action (or inaction). Conclusion This study highlights the complex interplay between physician actions and factors influencing physician-level variation in LoS. Interventions that target physicians but do not attend to team and hospital factors are likely insufficient to achieve sustained improvements in LoS. Aligning individual-level feedback and environmental restructuring with organizational values and needs of the patient population may offer a more promising approach to sustained improvement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.041 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.003 | 0.014 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".