Predicting emergency department use and unplanned hospitalization in patients with head and neck cancer: Development and validation of a machine learning algorithm.
Bibliographic record
Abstract
6021 Background: We recently demonstrated that patient reported symptom burden is strongly associated with emergency department use and unplanned hospitalization (ED/Hosp) in head and neck cancer (Noel et. al 2021 J. Clin Oncol - DOI: 10.1200/JCO.20.01845). We hypothesized that symptom scores could be used to build a tool that would accurately risk stratify patients. Methods: This was a population-based study of patients diagnosed with head and neck cancer between 2007 and 2018. All outpatient clinical encounters were identified. Edmonton Symptom Assessment Scores (ESAS) and clinical and demographic factors were abstracted. Training and test cohorts were randomly generated in a 4:1 ratio. Various machine learning algorithms were explored including: (1) logistic regression, (2) random forest, (3) gradient boosting machines (4) k-nearest neighbors and an (5) artificial neural network. Our main outcome was any 14-day ED/Hosp event following symptom assessment. The performance of each risk model was assessed on the test cohort using the area under the receiver operator characteristic (AUROC) curve and calibration plots. Shapley values were used to identify the variables with greatest contribution to the model. Results: The training cohort consisted of 9,409 patients undergoing 59,089 symptom assessments (80%). The remaining 2,352 patients and 14,193 symptom assessments were set aside as the test cohort (20%). Several models had high predictive accuracy, particularly the gradient boosting machine algorithm (validation AUROC 0.80 [95%CI 0.78-0.81]). A Youden-based cut-off corresponded to a validation sensitivity of 0.77 and specificity of 0.66. A second model built only with symptom severity data had an AUROC of 0.72 [95%CI 0.70-0.74]. Conclusions: Machine learning approaches can be used to predict ED/Hosp in head and neck cancer patients. This tool can risk stratify patients and may help direct targeted intervention.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".