Dynamic prediction of cancer-associated thrombosis to guide prophylactic anticoagulation.
Bibliographic record
Abstract
e13696 Background: Cancer-associated thrombosis (CAT) is preventable among high-risk individuals through prophylactic anticoagulation. Current guidelines recommend prophylactic anticoagulation based on risk assessment tools only applicable at the start of cancer treatment. However, patients face varying risks throughout their cancer journey. To address this, we developed a machine learning system to predict CAT risk longitudinally throughout cancer treatment, and evaluated anticoagulation prescription strategies using novel metrics of potential clinical utility. Methods: Using electronic health record data at Princess Margaret Cancer Centre, we assembled a cohort of thoracic and gastrointestinal cancer patients receiving systemic treatments between August 1, 2017 and December 31, 2019. We prompted open-source large language models (LLMs) to automatically detect CAT in 43,846 CT scan and doppler ultrasound reports, and manually labelled 5,000 CT scans and 350 dopplers for ground truth comparisons. Next, we trained longitudinal machine learning systems to predict CAT within 90 days of each cancer treatment. Models were tested on a held-out cohort of patients whose first treatment occurred in 2019. To assess clinical utility, we compared current guideline care, providing anticoagulation indefinitely to patients with a pre-treatment Khorana score > = 2, versus system-guided care, where patients would start anticoagulation whenever risk exceeds a threshold and stop after the risk remains below the threshold twice consecutively. We evaluated strategies based on the proportion of CAT potentially prevented, defined as anticoagulation recommended at least 14 days before CAT, against the average number of days per patient recommended for anticoagulation. Results: The overall cohort included 1,620 patients and 14,680 treatment sessions. When classifying radiology reports, the LLM (Mistral7B-Instruct) achieved an F1 score of 0.88 (95% CI, 0.81-0.92). CAT occurred within 90 days after 3.75% of treatment sessions. The system predicted the risk of CAT within 90 days with an area under receiver operating characteristic curve of 0.711 (95% CI, 0.663-0.757), outperforming the Khorana score. Across all risk thresholds, compared to the Khorana score, the system would potentially prevent more CAT, or require fewer average days of treatment. For example, when recommending the same average of 71.4 days on anticoagulation as the Khorana score, our system would increase the proportion of potentially prevented CAT by 10.8% (from 45.9% to 56.7%, P < 0.001). Conclusions: Machine learning can longitudinally predict CAT among patients receiving systemic therapy for cancer, outperforming existing approaches. These results show how personalized, longitudinal, machine learning guidance could prevent more CAT with fewer days on anticoagulation, enhancing effectiveness while reducing side effects and costs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".