P164 USING EXOME SEQUENCING TO EXPAND THE GENETIC ARCHITECTURE OF INFLAMMATORY BOWEL DISEASE
Bibliographic record
Abstract
IBD studies over the past decade have confirmed association to 250 gene loci. A handful of associations have led to specific validated functional variants highlighting intracellular response to microbes and regulation of adaptive immunity in IBD pathogenesis. For the vast majority, however, the specific implicated gene and causal functional variants remain unknown. In partnership with IBD researchers around the world, we launched an exome sequencing initiative to define the full allelic spectrum of protein-altering variation in genes associated to CD and/or UC, assess their role in clinical course and response to therapy, and to determine whether truncating variants confer risk or protection in each IBD gene in order to highlight opportune therapeutic targets. We have already completed exomes of 13,000 IBD cases providing a high-resolution view of coding variation at each GWAS hit and demonstrated a convincing excess of rare exome signal collectively. The cases are drawn from individual substudies that focus on isolated populations (Ashkenazi, Finnish, French-Canadian), admixed populations and clinical extremes, each providing unique opportunities for discovery. Early findings include novel protective truncating variants, the complete allelic series (including unique founder population alleles), non-additive inheritance models at known genes such as NOD2, overlooked low-frequency coding variants that explain GWAS hits (PRDM1, NRIP1). The integration of low frequency and rare functional variants with GWAS is moving us much closer towards completion of the genetic architecture of IBD. Early analyses of the data have revealed novel non-coding associations in known IBD genes as well as coding variants that implicate genes at loci in which the ‘causative’ gene was unknown. Furthermore, initial results have suggested variation in enrichment of different processes across different ethnicities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".