MétaCan
Menu
← Back to cohort

Dynamic prediction of cancer-associated thrombosis to guide prophylactic anticoagulation.

2025· article· en· W4410803412 on OpenAlexafffund
Jiang Chen He, Hosam Alghamni, Ian Hirsch, Baijiang Yuan, Muammar Kabir, Geoffrey Liu, Melanie Powis, Erik Yeo, Peter Groß, Benjamin Grant, Sharon Narine, Mattea Welch, Tran Truong, Robert C. Grant

Bibliographic record

VenueJournal of Clinical Oncology · 2025
Typearticle
Languageen
FieldMedicine
TopicVenous Thromboembolism Diagnosis and Management
Canadian institutionsUniversity of TorontoPrincess Margaret Cancer CentreUniversity Health Network
FundersPrincess Margaret Cancer Foundation
KeywordsMedicineThrombosisCancerIntensive care medicineVenous thromboembolismInternal medicineOncology

Abstract

fetched live from OpenAlex

e13696 Background: Cancer-associated thrombosis (CAT) is preventable among high-risk individuals through prophylactic anticoagulation. Current guidelines recommend prophylactic anticoagulation based on risk assessment tools only applicable at the start of cancer treatment. However, patients face varying risks throughout their cancer journey. To address this, we developed a machine learning system to predict CAT risk longitudinally throughout cancer treatment, and evaluated anticoagulation prescription strategies using novel metrics of potential clinical utility. Methods: Using electronic health record data at Princess Margaret Cancer Centre, we assembled a cohort of thoracic and gastrointestinal cancer patients receiving systemic treatments between August 1, 2017 and December 31, 2019. We prompted open-source large language models (LLMs) to automatically detect CAT in 43,846 CT scan and doppler ultrasound reports, and manually labelled 5,000 CT scans and 350 dopplers for ground truth comparisons. Next, we trained longitudinal machine learning systems to predict CAT within 90 days of each cancer treatment. Models were tested on a held-out cohort of patients whose first treatment occurred in 2019. To assess clinical utility, we compared current guideline care, providing anticoagulation indefinitely to patients with a pre-treatment Khorana score > = 2, versus system-guided care, where patients would start anticoagulation whenever risk exceeds a threshold and stop after the risk remains below the threshold twice consecutively. We evaluated strategies based on the proportion of CAT potentially prevented, defined as anticoagulation recommended at least 14 days before CAT, against the average number of days per patient recommended for anticoagulation. Results: The overall cohort included 1,620 patients and 14,680 treatment sessions. When classifying radiology reports, the LLM (Mistral7B-Instruct) achieved an F1 score of 0.88 (95% CI, 0.81-0.92). CAT occurred within 90 days after 3.75% of treatment sessions. The system predicted the risk of CAT within 90 days with an area under receiver operating characteristic curve of 0.711 (95% CI, 0.663-0.757), outperforming the Khorana score. Across all risk thresholds, compared to the Khorana score, the system would potentially prevent more CAT, or require fewer average days of treatment. For example, when recommending the same average of 71.4 days on anticoagulation as the Khorana score, our system would increase the proportion of potentially prevented CAT by 10.8% (from 45.9% to 56.7%, P < 0.001). Conclusions: Machine learning can longitudinally predict CAT among patients receiving systemic therapy for cancer, outperforming existing approaches. These results show how personalized, longitudinal, machine learning guidance could prevent more CAT with fewer days on anticoagulation, enhancing effectiveness while reducing side effects and costs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.016
Threshold uncertainty score0.033

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.012
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0020.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.000
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.086
GPT teacher head0.491
Teacher spread0.405 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueJournal of Clinical Oncology→Same topicVenous Thromboembolism Diagnosis and Management→French-language works237,207→