Non-linear Mendelian randomization: evaluation of effect modification in the residual and doubly-ranked methods with simulated and empirical examples
Bibliographic record
Abstract
Non-linear Mendelian randomisation (NLMR) is a relatively recently developed approach to estimate the causal effect of an exposure on an outcome where this is expected to be non-linear. Two commonly used techniques-based on stratifying the exposure and performing Mendelian randomisation (MR) within each strata-are the residual and doubly-ranked methods. The residual method is known to be biased in the presence of genetic effect heterogeneity-where the effect of the genotype on the exposure varies between individuals. The doubly-ranked method is considered to be less sensitive to genetic effect heterogeneity. In this paper, we simulate genetic effect heterogeneity and confounding of the exposure and outcome and identify that both methods are susceptible to likely unpredictable bias in this setting. Using UK Biobank, we identify empirical evidence of genetic effect heterogeneity and show via simulated outcomes that this leads to biased MR estimates within strata, whilst conventional MR across the full sample remains unbiased. We suggest that these biases are highly likely to be present in other empirical NLMR analyses using these methods and urge caution in current usage. Simulated outcome analyses may represent a useful test to identify if genetic effect heterogeneity is likely to bias NLMR estimates in future analyses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".