Continuous Covariates in Genetic Association Studies of Case-Parent Triads: Gene and Gene-Environment Interaction Effects, Population Stratification, and Power Analysis
Bibliographic record
Abstract
We propose a multinomial logistic regression method which permits estimation and likelihood ratio tests for allele effects, their interactions with continuous covariates, and assessment of the degree of population stratification in genetic association studies of case-parent triads. Our approach overcomes the constraint imposed by the categorical nature of explanatory variables in the log-linear model. We also demonstrate that the multinomial logistic method can yield efficient inference in the presence of missing parental genotype data via the use of the Expectation-Maximization (EM) algorithm. We performed simulations to compare the multinomial logistic model with the case-pseudosibling conditional logistic model approach, both of which permit the incorporation of continuous covariates. Simulation results indicate that the multinomial logistic model and the conditional logistic model lead to similar estimates in large samples. A simulation-based method of sample size estimation is also used to show that the two models are approximately equivalent in sample size requirements. When parental genotype data are missing, either completely at random or dependent on covariates, the use of the EM algorithm gives multinomial logistic model greater power. Since the multinomial logistic model offers the possibility of assessing the degree of population stratification in the sample and can also provide efficient inference in the presence of missing parental genotypes, the proposed model has an important application in epidemiological family-based association studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".