On the Use of Population Attributable Fraction to Determine Sample Size for Case-Control Studies of Gene-Environment Interaction
Bibliographic record
Abstract
Most methods for calculating the sample size needed to detect gene-environment interactions use odds ratios to measure the effect size. We show that for any combination of susceptible genotype prevalence and exposure prevalence and their associated risks, the odds ratio measuring strength of interaction corresponds to a population attributable fraction (PAF) because of interaction and vice versa. Simultaneous consideration of odds ratio for interaction and the associated PAF attributable to interaction provides additional insight to investigators evaluating the feasibility and public health relevance of a proposed study. We considered gene-environment interactions on a multiplicative scale, and assumed a dichotomous environmental exposure variable and a single two-allele disease-susceptibility locus. Our results show, for example, that for studies of exposures and genotypes that are common in a population (30%-50%), the PAF for interaction is large (>27%) even if the odds ratio for interaction is only moderate (approximately 2). If simultaneous estimates of interaction odds ratio and PAF indicate that the PAF is so large as to be implausible, the investigator may decide to reevaluate the study design based on detecting a more reasonable PAF. In this case, the associated odds ratio for interaction will be weaker and a considerably larger sample size may be needed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.052 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".