Candidate Gene Association Resource (CARe)
Bibliographic record
Abstract
BACKGROUND: The National Heart, Lung, and Blood Institute's Candidate Gene Association Resource (CARe), a planned cross-cohort analysis of genetic variation in cardiovascular, pulmonary, hematologic, and sleep-related traits, comprises >40,000 participants representing 4 ethnic groups in 9 community-based cohorts. The goals of CARe include the discovery of new variants associated with traits using a candidate gene approach and the discovery of new variants using the genome-wide association mapping approach specifically in African Americans. METHODS AND RESULTS: CARe has assembled DNA samples for >40,000 individuals self-identified as European American, African American, Hispanic, or Chinese American, with accompanying data on hundreds of phenotypes that have been standardized and deposited in the CARe Phenotype Database. All participants were genotyped for 7 single-nucleotide polymorphisms (SNPs) selected based on prior association evidence. We performed association analyses relating each of these SNPs to lipid traits, stratified by sex and ethnicity, and adjusted for age and age squared. In at least 2 of the ethnic groups, SNPs near CETP, LIPC, and LPL strongly replicated for association with high-density lipoprotein cholesterol concentrations, PCSK9 with low-density lipoprotein cholesterol levels, and LPL and APOA5 with serum triglycerides. Notably, some SNPs showed varying effect sizes and significance of association in different ethnic groups. CONCLUSIONS: The CARe Pilot Study validates the operational framework for phenotype collection, SNP genotyping, and analytic pipeline of the CARe project and validates the planned candidate gene study of approximately 2000 biological candidate loci in all participants and genome-wide association study in approximately 8000 African American participants. CARe will serve as a valuable resource for the scientific community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.045 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.007 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.008 | 0.005 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.284 | 0.119 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".