Machine Learning–Based Early Warning Systems for Acute Care Utilization During Systemic Therapy for Cancer
Bibliographic record
Abstract
BACKGROUND: Emergency department visits and hospitalizations frequently occur during systemic therapy for cancer. We developed and evaluated a longitudinal warning system for acute care use. METHODS: Using a retrospective population-based cohort of patients who started intravenous systemic therapy for nonhematologic cancers between July 1, 2014, and June 30, 2020, we randomly separated patients into cohorts for model training, hyperparameter tuning and model selection, and system testing. Predictive features included static features, such as demographics, cancer type, and treatment regimens, and dynamic features, such as patient-reported symptoms and laboratory values. The longitudinal warning system predicted the probability of acute care utilization within 30 days after each treatment session. Machine learning systems were developed in the training and tuning cohorts and evaluated in the testing cohort. Sensitivity analyses considered feature importance, other acute care endpoints, and performance within subgroups. RESULTS: The cohort included 105,129 patients who received 1,216,385 treatment sessions. Acute care followed 182,444 (15.0%) treatments within 30 days. The ensemble model achieved an area under the receiver operating characteristic curve of 0.742 (95% CI, 0.739-0.745) and was well calibrated in the test cohort. Important predictive features included prior acute care use, treatment regimen, and laboratory tests. If the system was set to alarm approximately once every 15 treatments, 25.5% of acute care events would be preceded by an alarm, and 47.4% of patients would experience acute care after an alarm. The system underestimated risk for some treatment regimens and potentially underserved populations such as females and non-English speakers. CONCLUSIONS: Machine learning warning systems can detect patients at risk for acute care utilization, which can aid in preventive intervention and facilitate tailored treatment. Future research should address potential biases and prospectively evaluate impact after system deployment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".