MétaCan
Menu
Back to cohort
Record W4415882029 · doi:10.1101/2025.11.02.25339322

Target trial emulation of physical activity and cardiovascular disease risk: What is impact of the exposure assessment method?

2025· preprint· W4415882029 on OpenAlexaff
Matthew Ahmadi, Borja del Pozo Cruz, Raaj Kishore Biswas, Armando Teixeira-Pinto, Marla Beauchamp, Joanne McVeigh, Mark Hamer, Emmanuel Stamatakis

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Language
FieldMathematics
TopicAdvanced Causal Inference Techniques
Canadian institutionsMcMaster University
Fundersnot available
KeywordsCausal inferenceConfoundingObservational studyRandomized controlled trialPedometerRisk assessmentCohortPhysical activityPoisson regression

Abstract

fetched live from OpenAlex

Background: Target trial emulation (TTE) designs provide a framework for strengthening causal inference in observational research, but it is unknown how vulnerable they are to substantial error when the emulated intervention (exposure) measurement is imprecise. In physical activity epidemiology specifically, correcting for confounding via TTE designs but not addressing the large measurement errors arising from self-reports (e.g. questionnaires, which typically capture partial behavioural accounts with very low precision) creates uncertainty about possible dominance of type 2 error biases arising from such novel designs. High-resolution wearables-based methods capture most movement, providing an assessment of physical activity behaviour with substantially less, empirically verified, measurement error. No study has examined how physical activity measurement method influences causal inference in TTE studies. Objectives: We applied TTE methodology to sub-samples of the UK Biobank cohort with repeat exposure measurements, to compare the effects of an emulated physical activity intervention on incident CVD risk, when the physical activity exposure was quantified using self-report vs. wearable devices. Methods: The emulated randomized controlled trial identified physically inactive adults (<150 moderate-to-vigorous physical activity (MVPA) mins/week) who had repeat assessments for wearable and self-reported physical activity. At re-examination, participants were categorised into intervention (adopted the current recommendation of ≥150 MVPA mins/week) or control (remained physically inactive) groups. Participants in each group were propensity score-matched to balance lifestyle behaviours, demographic, and health factors. Cumulative risk for CVD incidence was assessed through cumulative risk curves, hazard ratios, risk ratios, using Fine-Gray subdistribution and Poisson regression models. Results: The wearables analytic sample included 490 participants (245 per arm; mean incident CVD follow-up 4.4 years), and the self-report sample included 11,302 participants (5,651 per arm; mean follow-up 6.3 years). In wearables assessments, guideline-adherent participants had markedly lower cumulative CVD risk (cumulative risk = 8.0% vs. 17.0%; hazard ratio [95%CI] = 0.59 [0.36, 0.98]; relative risk = 0.45 [0.28, 0.72]). In contrast, self-report assessments showed near-identical risk trajectories for intervention and control groups (cumulative risk = 21.6% vs. 21.2%; hazard ratio = 0.98 [0.89, 1.08]; relative risk = 0.92 [0.84, 1.00]). Matching the self-report sample to the wearables sample for lifestyle, demographic, and health factors confirmed these findings. Conclusion: Reliance on self-reported measures of physical activity in TTE studies may obscure emulated intervention effects due to non-differential misclassification, increasing considerably risk of Type II error. Exposure assessment using wearable devices may be essential for valid causal inference in TTE studies of physical activity and CVD risk. Future TTE studies of physical activity exposures should prioritise objective measurements to avoid biased inferences that could affect public health policy and guidelines.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.532
metaresearch head score (Gemma)0.738
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.468
Threshold uncertainty score0.578

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5320.738
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0090.016
Bibliometrics0.0020.003
Science and technology studies0.0010.006
Scholarly communication0.0060.009
Open science0.0060.006
Research integrity0.0080.008
Insufficient payload (model declined to judge)0.0090.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.083
GPT teacher head0.443
Teacher spread0.360 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicAdvanced Causal Inference TechniquesFrench-language works237,207