Development of patient reported outcomes-based machine learning algorithm for the six-month mortality prediction in patients with advanced cancer.
Bibliographic record
Abstract
273 Background: To date, studies of machine learning (ML) algorithms within oncology for mortality prediction have focused on structured electronic health record (EHR) data. Given the complex symptom burden of patients with advanced cancers, ML models may be better suited to identify patterns and interactions between symptom burden and outcomes compared to traditional statistical methods. To that end, in this study, we leverage the patient reported outcomes (PRO) data together with clinical EHR-based variables to assess the performance of ML algorithms to predict mortality in patients with advanced cancers. Methods: We randomly selected 689 patients with advanced cancer who had their first Palliative Care encounter between January 2012 and December 2017. 59 patients were lost to follow-up and were excluded from this analysis. The remaining cohort of 630 patients was split 4:1 randomly into a training and validation set to develop and test a supervised ML algorithm (Extreme Gradient Boosting [XGB] tree) to predict the 6-month mortality. Candidate variables for algorithm development included gender, age, ECOG performance status (PS), number of prior systemic therapies, and scores on the Edmonton Symptom Assessment System (ESAS)-FS, a 12-item PRO measure of physical and psychosocial symptom burden include the composite Physical Symptom Score (PHS), a sum of the physical ESAS symptoms (pain, fatigue, nausea, drowsiness, shortness of breath, appetite, wellbeing, sleep). Results: Overall, 630 patients were included in this 6-month mortality prediction; mean age 59 years, 354 (56%) female; 276 (44%) male. Variables with the most significant impact on the XGB tree mortality prediction were the ESAS symptoms of shortness of breath (1-AUC, 0.295), appetite, ESAS PHS, financial distress, age, and appetite as well as ECOG PS and number of prior systemic therapies. The XGB tree algorithm demonstrated the best overall prediction performance of 6-month mortality in the independent testing set, AUC 0.716 (95% CI 0.63 - 0.81), sensitivity 0.75 (95% CI 0.66 - 0.87), and a positive predictive value 0.67 (95% CI 0.57 - 0.79). Conclusions: Our ML model leveraged PRO-based assessment of symptom burden to correctly identify the majority of patients who died within 6 months. These models are uniquely positioned to not only automatically identify patients at high risk for short-term mortality but also the specific symptoms of concern for clinical intervention. Such models can be applied to available clinical and PRO data to facilitate clinical decision-making. Futures studies on improving model performance with the inclusion of interventions to modify symptom burden are in design.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".