Deep resequencing identifies candidate functional genes in leprosy GWAS loci
Bibliographic record
Abstract
Leprosy is the second most prevalent mycobacterial disease globally. Despite the existence of an effective therapy, leprosy incidence has consistently remained above 200,000 cases per year since 2010. Numerous host genetic factors have been identified for leprosy that contribute to the persistently high case numbers. In the past decade, genetic epidemiology approaches, including genome-wide association studies (GWAS), identified more than 30 loci contributing to leprosy susceptibility. However, GWAS loci commonly encompass multiple genes, which poses a challenge to define causal candidates for each locus. To address this problem, we hypothesized that genes contributing to leprosy susceptibility differ in their frequencies of rare protein-altering variants between cases and controls. Using deep resequencing we assessed protein-coding variants for 34 genes located in GWAS or linkage loci in 555 Vietnamese leprosy cases and 500 healthy controls. We observed 234 nonsynonymous mutations in the targeted genes. A significant depletion of protein-altering variants was detected for the IL18R1 and BCL10 genes in leprosy cases. The IL18R1 gene is clustered with IL18RAP and IL1RL1 in the leprosy GWAS locus on chromosome 2q12.1. Moreover, in a recent GWAS we identified an HLA-independent signal of association with leprosy on chromosome 6p21. Here, we report amino acid changes in the CDSN and PSORS1C2 genes depleted in leprosy cases, indicating them as candidate genes in the chromosome 6p21 locus. Our results show that deep resequencing can identify leprosy candidate susceptibility genes that had been missed by classic linkage and association approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".