Genomic analysis of Canadian children with food allergies points to the immunoglobulin heavy chain gene locus
Notice bibliographique
Résumé
Analysis of a pediatric food allergy population using whole-genome sequencing points to the immunoglobulin heavy chain gene locus. This region is critical for immune responses, as rearrangements of its variable, diversity, and joining domains generate a highly diversified antibody repertoire. However, because it is highly variable and repetitive, it can only be reliably studied using sequencing approaches, leaving its genetic variants poorly documented. Food allergy is a major public health concern, affecting about 8% of children in Western countries, 40% of whom are polyallergic.1 The standard approach to managing food allergy is complete avoidance of food allergens. However, avoidance is complicated by the broad presence of the most common allergens. Since symptoms vary from mild to life-threatening anaphylactic shock,2 it is crucial to better understand the biology underlying the development, diversification, and severity of food allergy. To address this need, a pediatric research clinic was established in Saguenay–Lac-Saint-Jean, a region in northeastern Quebec, Canada. The aim of this initiative is to facilitate access to oral immunotherapy, an emerging treatment for food allergy, and to develop the Zéro allergie cohort.3 This cohort was designed to study the genomic, epigenomic, metabolomic, and microbial diversity associated with food allergy and oral immunotherapy outcomes. This study aimed to analyze the genomic profile of 100 children from the Zéro allergie cohort to identify genetic variants associated with food allergy in a Canadian pediatric population. Each allergic child recruited in the cohort was selected to undergo oral immunotherapy at the research clinic for one to three food allergens (Table 1). Diagnoses of FA were confirmed by a pediatrician based on clinical symptoms and skin prick test (wheal diameter ≥3 mm of the negative control, measured after 10 min; see Table S1 for the list of allergens represented in each category). Non-allergic siblings were also recruited as controls. Children of all sexes, ages (from 6 months to 17 years), and ethnicities were included. See Appendix S1 for details on recruitment. In the present analyses, all children were of Caucasian ethnicity. Significant differences between the food allergy and control groups were observed for the median age (3.57 ± 2.54 years vs. 5.15 ± 2.62 years, p = .027), mean birth order (1.82 ± 0.84 vs. 1.33 ± 0.49, p = .017), the number of children who were breastfed (69 children (86%) vs. 11 children (61%), p = .020), duration of breastfeeding (31.66 ± 24.80 weeks vs. 21.83 ± 33.65 weeks, p = .026), and the prevalence of atopic dermatitis (69 children (86%) vs. 10 children (56%), p = .006). Sex, number of siblings at home, prevalence of preterm birth, caesarean section, and asthma were similar between the two groups (p > .05). The number of past or present food allergies among allergic children ranged from one to seven, with a median of two. The number of food allergens for which children were desensitized ranged from one to four, with a median of one. One child underwent desensitization for more than three allergens because a fourth allergy was identified during the desensitization protocol. In the control group, one child had a resolved egg allergy, and another had a resolved milk allergy, resulting in the number of past or present food allergies among controls ranging from zero to one (median = zero). After quality control, whole-genome sequencing data were available for 80 allergic children and 18 non-allergic siblings (see Appendix S1 for the complete method description). Among the phenotypic variables that differed significantly between groups, age at sampling, birth order, and breastfeeding status were included in the analyses as covariates. Duration of breastfeeding was not considered due to its strong association with the binary breastfeeding phenotype (p-value = 3.32 × 10−11), in order to avoid redundancy in statistical models. The proportion of atopic dermatitis was not included in the analysis models because of its strong link to the prevalence of food allergy through the atopic march concept.4 A first analysis was conducted for food allergy, regardless of the specific food allergen, including the first two principal components as fixed covariates and a matrix of estimated relatedness as a random effect (see Appendix S1 for details on analyses). No association reached the genome-wide significance threshold of 5 × 10−8, but suggestive associations (p < 1 × 10−5) were identified for 100 variants (Table S2). Compared with the analyses that included the additional selected covariates, four loci were found in both analyses (9q33.3 from positions 124,448,987 to 124,450,621; 12q24.32–24.33 from positions 127,816,100 to 128,765,615; 14q32.33 from positions 106,414,608 to 106,419,471; and 16p13.3 from positions 2,789,758 to 2,916,358, Figure 1A,B and Table S3). Among these, locus 14q32.33 emerged as the most consistent association across the two models, with 16 genetic variants in common and the strongest association signals overall when the two analyses were considered together (p = 8.98 × 10−6–1.23 × 10−7 in the first model; p = 9.37 × 10−6–7.51 × 10−7 in the fully adjusted model). This locus lies within a non-coding region of the immunoglobulin heavy (IGH) chain gene cluster and is surrounded by functional and non-functional sequences of immunoglobulin segments for the variable domain (V domain; Figure 1E). These segments play a key role in immunity by rearranging with diversity and joining domains (D and J domains) to generate highly diverse immunoglobulins in developing B cells.5, 6 They are essential for producing a broad repertoire of B cell and T cell receptors to mount an appropriate immune response.6 Analysis of antibody repertoire from transcriptomic data showed that B cells from children with food allergen sensitization exhibit skewed V-domain usage, potentially reflecting constitutional differences in recombination mechanisms in food-allergic individuals.7 Sequencing of antibody repertoires has led to the identification of coding variants and rearrangements, but the locus as a whole remains poorly documented, as its study requires DNA sequencing data. In fact, only four variants in the IGH chain locus are represented on the Illumina Infinium ImmunoArray24v2, and imputation using reference sequences is hampered by the high variability and number of repeats in this region.6 To date, more than 680 IGH alleles are catalogued from antibody repertoires analyses.5, 8 A recent study using both long-read and antibody repertoire sequencing identified, within the IGH locus, 20,510 single nucleotide variants, 966 indels, 71 small structural variants, and eight large structural variants (>9 Kbp) from only 154 participants,5 highlighting the extreme variability of this region. The positions identified in the present study fall at the beginning of one of these large structural variants,5 although there is currently no evidence of its possible impact on rearrangement mechanisms. Since most previous genome-wide association studies have relied on microarray and imputation data, no other genetic associations with food allergy have been reported to date. As children who experienced transient food allergy may share the same genetic profile as other food allergic children, analyses were also conducted considering these children as allergic rather than as controls (Figure 1C,D, Tables S4 and S5). Notably, the only locus validated by those analyses was 14q32.33. To assess possible specificity of the signal according to allergen category, exploratory analyses were carried out, indicating an association with the IGH locus in the 14q32.33 region for past or present egg or milk allergies (n = 29), with a significant signal observed at positions 105,661,635 to 105,774,118 (Figure S1A–D). These specific positions were also identified through the association of a variant in the original analysis (position 105,770,197) and are located at the beginning of the IGH locus (Figure S1E). However, the lack of significant signals in the different allergen categories for the region covering the variable domain genes of this locus may be attributable to the small number of affected individuals per category or to an additive effect of all food allergy types in the initial analysis. The specificity of the results with respect to other allergic diseases, such as asthma and eczema (atopic dermatitis), was also evaluated, and none showed an association with the 14q32.33 locus (Figure S2A,B). Even though their link with food allergy is more indirect, the three other loci are also of interest as they point to genes involved in immune response or asthma, another disease within the atopic march.4 The 9q33.3 locus is near the ADGRD2 gene, which was previously associated with white blood cell and neutrophil counts.9 The 12q24.32–24.33 and 16p13.3 loci span the TMEM132C and FLYWCH1 genes, respectively, both of which have been linked to forced expiratory volume in 1 s,10, 11 a proxy of asthma phenotype. This study has some limitations. The relatively small and imbalanced number of controls reduces statistical power, although this is preferable to an equal distribution with a smaller total sample size. The imbalance reflects the difficulty of recruiting non-allergic siblings as some families include only allergic children and of obtaining parental consent for invasive blood sampling. Another limitation arises from the high variability and repetitive nature of the immunoglobulin regions, for which long-read sequencing would capture more variants.6 Still, short-read sequencing is more accurate for these regions than microarrays and imputation, and it enables the identification of an allergy-related locus highly relevant to food allergy. Recruitment for the Zero allergy cohort is ongoing. It will allow sequencing of additional samples from allergic and non-allergic siblings, as well as long-read sequencing to further analyze the immunoglobulin heavy and light chain loci, including the impact of associated non-coding variants on the complex rearrangements within these regions. This study uncovered four loci associated with food allergy, most notably within the highly variable IGH chain gene region at chromosome 14q32.33. Despite its likely causal relevance, this locus remains largely unexplored due to its repetitive and complex structure. A better understanding of this region in the context of food allergy could reveal key genetic mechanisms underlying both risk and protective factors, ultimately enabling improved risk prediction and targeted interventions. Anne-Marie Madore: Formal analysis; data curation; methodology; writing – original draft; writing – review and editing. Marie-Ève Lavoie: Methodology; writing – review and editing. Anne-Marie Boucher-Lafleur: Methodology; writing – review and editing. Frédérique Gagnon-Brassard: Methodology; writing – review and editing. Noémie Gilbert: Methodology; writing – review and editing. Philippe Bégin: Methodology; validation; writing – review and editing. Sarah Lavoie: Methodology; writing – review and editing. Cloé Rochefort-Beaudoin: Methodology; writing – review and editing. Claudia Nuncio-Naud: Methodology; writing – review and editing. Guy Parizeault: Methodology; writing – review and editing. Charles Morin: Methodology; writing – review and editing. Catherine Laprise: Project administration; supervision; resources; conceptualization; funding acquisition, writing – review and editing. The authors are grateful to all the children and their families of the Zéro allergie cohort who participated in this study. This work was supported by the Canada Research Chair in Genomics of Asthma and Allergic Diseases (award number CRC-2021-00108), Fondation de l'Université du Québec à Chicoutimi (https://fuqac.ca/) and Fondation de ma vie (https://www.fondationdemavie.qc.ca/). The authors declare no conflict of interest in relation to this study. Informed consent was obtained from the legal guardians of all participants. The experimental protocol was approved by the Ethics Committee of the Centre intégré universitaire de santé et de services sociaux du Saguenay–Lac-Saint-Jean (project #2022–015). Appendix S1. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,003 | 0,004 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».