A mathematical model and inference method for bacterial colonization in hospital units applied to active surveillance data for carbapenem-resistant enterobacteriaceae
Bibliographic record
Abstract
Widespread use of antibiotics has resulted in an increase in antimicrobial-resistant microorganisms. Although not all bacterial contact results in infection, patients can become asymptomatically colonized, increasing the risk of infection and pathogen transmission. Consequently, many institutions have begun active surveillance, but in non-research settings, the resulting data are often incomplete and may include non-random testing, making conventional epidemiological analysis problematic. We describe a mathematical model and inference method for in-hospital bacterial colonization and transmission of carbapenem-resistant Enterobacteriaceae that is tailored for analysis of active surveillance data with incomplete observations. The model and inference method make use of the full detailed state of the hospital unit, which takes into account the colonization status of each individual in the unit and not only the number of colonized patients at any given time. The inference method computes the exact likelihood of all possible histories consistent with partial observations (despite the exponential increase in possible states that can make likelihood calculation intractable for large hospital units), includes techniques to improve computational efficiency, is tested by computer simulation, and is applied to active surveillance data from a 13-bed rehabilitation unit in New York City. The inference method for exact likelihood calculation is applicable to other Markov models incorporating incomplete observations. The parameters that we identify are the patient-patient transmission rate, pre-existing colonization probability, and prior-to-new-patient transmission probability. Besides identifying the parameters, we predict the effects on the total prevalence (0.07 of the total colonized patient-days) of changing the parameters and estimate the increase in total prevalence attributable to patient-patient transmission (0.02) above the baseline pre-existing colonization (0.05). Simulations with a colonized versus uncolonized long-stay patient had 44% higher total prevalence, suggesting that the long-stay patient may have been a reservoir of transmission. High-priority interventions may include isolation of incoming colonized patients and repeated screening of long-stay patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".