Confidence intervals for candidate gene effects and environmental factors in population‐based association studies of families
Bibliographic record
Abstract
Complex diseases are influenced by both genetic and environmental factors. Studies of individuals or of families can be used to examine the association of genetic factors, such as candidate genes, and other risk factors with the presence or absence of complex disorders. If families are investigated, whether or not they are randomly ascertained, possible familial correlation among observations must be considered. We have compared two statistical approaches for analyzing correlated binary data from randomly ascertained nuclear families. The generalized estimating equations approach (GEE) can be used to adjust for familial correlation. The relationship between covariates and the response is modelled, and the correlations among family members are treated as nuisance parameters. For comparison, we have proposed two strategies from a hierarchical nonparametric bootstrap approach. One strategy (S1) samples family units, preserving the structure and correlation within each family. A second and novel strategy (S2) also samples family units but then randomly samples offspring with replacement in each family. We applied the methods to data from a study of cardiovascular disease, and followed up with a simulation study in which family data were generated from an underlying multifactorial genetic model. Although the bootstrap approach was more computationally demanding, it outperformed the GEE in terms of confidence interval coverage probabilities for all sample sizes considered.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".