MétaCan
Menu
Back to cohort
Record W4416222200 · doi:10.2196/76659

Estimating 10-Year Cardiovascular Disease Risk in Primary Prevention Using UK Electronic Health Records and a Hybrid Multitask BERT Model: Retrospective Cohort Study

2025· article· en· W4416222200 on OpenAlexvenueno aff
Tianyi Liu, Lei Lü, Yanzhong Wang, Andrew J. Krentz, Vasa Ćurčin

Bibliographic record

VenueJMIR Medical Informatics · 2025
Typearticle
Languageen
FieldComputer Science
TopicMachine Learning in Healthcare
Canadian institutionsnot available
Fundersnot available
KeywordsRetrospective cohort studyHealth recordsPrimary carePrimary preventionDiseaseRisk stratificationMEDLINECohort studyCohort

Abstract

fetched live from OpenAlex

Background: Cardiovascular disease (CVD) remains a leading cause of preventable morbidity and mortality, highlighting the need for early risk stratification in primary prevention. Traditional Cox models assume proportional hazards and linear effects, limiting flexibility. While machine learning offers greater expressiveness, many models rely solely on structured data and overlook time-to-event (TTE) information. Integrating structured and textual representations may enhance prediction and support equitable assessment across clinical subgroups. Objective: This study aims to develop a hybrid multitask deep learning model (MT-BERT [multitask Bidirectional Encoder Representations from Transformers]) integrating structured and textual features from electronic health records (EHRs) to predict 10-year CVD risk, enhancing individualized stratification and supporting equitable assessment across diverse demographic groups. Methods: We used data from Clinical Practice Research Datalink (CPRD) Aurum comprising 469,496 patients aged 40-85 years to develop MT-BERT for 10-year CVD risk prediction. Structured EHR variables and their corresponding textual representations were jointly encoded using a multilayer perceptron and a distilled version of the BERT model (DistilBERT), respectively. A fusion layer and stacked multihead attention modules enabled cross-modal interaction modeling. The model generated both binary classification outputs and TTE risk scores, optimized using a custom FocalCoxLoss function with uncertainty-based weighting. Prediction targets encompassed composite and individual CVD outcomes. Model performance was evaluated using the area under the receiver operating characteristic curve (AUROC), concordance index, and Brier score, with subgroup analyses by ethnicity and deprivation, and heterogeneity assessed using Higgins I² and Cochran Q statistics. Generalizability was assessed via external validation in a held-out London cohort. Results: The MT-BERT model yielded AUROC values of 0.744 (95% CI 0.738-0.749) in males and 0.782 (95% CI 0.768-0.796) in females on the test set (n=711,052), and 0.736 (95% CI 0.729-0.741) and 0.775 (95% CI 0.768-0.780), respectively in "spatial external" validation (n=144,370). Brier scores were 0.130 in males and 0.091 in females. Individuals classified as high-risk (≥40% risk in males and ≥34% in females) demonstrated significantly reduced 10-year event-free survival relative to lower-risk individuals (log-rank P<.001). Model performance was consistently higher in females across all metrics. Subgroup analyses revealed substantial heterogeneity across ethnicity and deprivation (I²>70%), especially among males, with lower AUROC in South Asian and Black ethnic groups. These findings reflect variation in model performance across demographic groups while supporting its applicability to large-scale CVD risk stratification. Conclusions: The proposed hybrid MT-BERT model predicts 10-year CVD risk for primary prevention by integrating structured variables and unstructured clinical text from EHRs. Its multitask design facilitates both individualized risk stratification and TTE estimation. While performance was modestly reduced in deprived and minority ethnic subgroups, these findings provide preliminary support for advancing equity-aware, data-driven prevention strategies in increasingly diverse health care settings.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.763
Threshold uncertainty score0.834

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.011
GPT teacher head0.318
Teacher spread0.307 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Medical InformaticsSame topicMachine Learning in HealthcareFrench-language works237,207