Establishing association between <i>HLA‐C*04:01</i> and severe <scp>COVID</scp>‐19
Bibliographic record
Abstract
HLA is a multi-gene locus essential to acquired immunity in humans.1 HLA molecules trigger immune responses to pathogen challenges by presenting antigen-derived epitopes on the cell surface to T cells, making them vital to human health and medicine. In this letter, we report on the growing body of evidence linking HLA-C*04:01 and severe COVID-19. Starting in late 2019, the severe acute respiratory syndrome (SARS)-CoV-2 coronavirus virus wreaked havoc on the world, and quickly the World Health Organization (WHO) declared the associated disease, COVID-19, a pandemic on March 11, 2020.2 While many people succumbed to the disease3 (https://coronavirus.jhu.edu/map.html), a great many also showed no symptoms of infection despite testing positive for the SARS-CoV-2 virus.4 Reports of variable frequency of asymptomatic individuals in certain jurisdictions hinted at a population-level immunity, with a potential genetic underpinning. Since HLA genes are present in human populations at observed frequencies5 (http://allelefrequencies.net/), it is the usual “molecular” suspect when certain demographics appear more or less at risk. Although protective HLA alleles have been observed previously in the context of pathogen-caused illnesses,6 there are stronger established links between HLA alleles and disease susceptibility and prognosis.7 Although the exact reason for this is unclear, it has been previously postulated that HLA alleles with a limited ability to bind antigen-derived epitopes may contribute to a deficient immune response since they would have less T cell visibility.8 This may be exploited by rapidly mutated viruses, for example; those whose successful particles evade the immune system and circumvent host defenses with stealth profiles that may allow them to proliferate unchecked, in some of their hosts. Early on in the COVID-19 pandemic, numerous research groups, including ours, analyzed rapidly generated patient data for links between patient HLA and COVID-19 susceptibility, and progression—including disease severity. Genome and/or transcriptome sequencing datasets coupled with detailed COVID-19 patient outcome clinical indicators were made publicly available as early as mid-2020,9 enabling the scientific community working together to counter this global threat. We had reported on our initial analysis of a heterogenous, albeit small (n = 126, with 26 non-COVID control), New-York COVID-19 patient cohort,9 identifying HLA alleles C*04:01 and to a lesser extent A*11:01, as disease severity markers—patients in need of mechanical ventilation.10 Statistical significance in this early analysis was marginal, and we had cautioned our readers about the interpretability of association results derived from a limited COVID-19 patient cohort size, advocating for larger patient cohorts from diverse communities.11 Since our initial report,10 several studies have recapitulated associations of HLA-C*04:01 and A*11:01 alleles with COVID-19 severity (reviewed in References 12, 13), including a validation of our findings on the New York cohort data,14 but also extending to patient cohorts across the globe including Europe,15 India,16 Armenia,17 Spain,18 and the United Arab Emirates19 (Figure 1). While the patient cohort sizes in most of these studies was largely on the same order of magnitude with that of Overmyer's (i.e., hundreds of patients), we note that recapitulating the C*04:01 association in these independent studies is not necessarily a given, especially when considering the sheer number of HLA class I alleles reported to date20 (>24,000). Also, C*04:01 allele frequency fluctuates widely in human populations5 (min–max = 3.8%–22.8%, median 13.0% in 60 population studies with sample size >1000; http://allelefrequencies.net/). In fact, it is worth pointing out that reports looking for HLA associations with COVID-19, including some analysis of large patient cohorts carrying genome-wide association studies (GWAS), have revealed either nothing of significance,21, 22 or associations between COVID-19 disease prognosis markers and HLA class I (and class II) alleles other than C*04:01.12, 13 It is possible that despite the prevalence of this allele in human populations, its frequency in those cohorts was insufficient to register significant associations with COVID-19 disease severity. Still, despite the inconsistence or absence of findings reported in those earlier reports, the C*04:01 and COVID-19 severity association has so far been independently recapitulated and reported in six independent, peer-reviewed studies, including within a large (n = 9300) Spanish COVID-19 patient cohort.18 It has also been re-observed by us independently in a larger and ethnically diverse cohort of Canadian COVID-19 patients (CanCOGeN CGEn HostSeq n = 9460, with 8328 (88.0%) COVID-19 positive patients23) (Figure 2). Although C*04:01 has a restrictive ability to bind only six SARS-CoV-2-derived epitopes—one of the worst predicted HLA binder to SARS-CoV-2-derived peptides17, 24, 25—our initial analysis of the CanCOGeN data does not reveal a preferential link between HLA alleles having a limited ability to bind SARS-CoV-2 epitopes when COVID-19, ICU and disease severity status are considered (Figure 2). Further, the other HLA-I alleles identified by our analysis (e.g., A*30:01 and B*53:01) had a low allele frequency in our sample (1.8% and 1.2%, respectively) compared with C*04:01 (13.2%) (Figure 2). It is worth noting that B*53:01 is a relatively low-frequency allele in the population5 (min–max = 0.01%–13.3%, median 0.6% in 77 population studies with sample size >1000; http://allelefrequencies.net/) and, partly for this reason, a link with severe COVID-19 may not be readily observed, with exceptions.26 Despite the low frequency of this allele in our sample (1.2%, n = 233), 76% haplotyped with C*04:01 and together, both alleles associated strongly with severe COVID-19 in the CanCOGeN cohort (Bonferroni-corrected FET p = 7.0 × 10−6). Even though the COVID-19 pandemic has now turned endemic,27, 28 many large-scale studies are still processing data collected during the pandemic. We expect the results from these studies will continue to yield valuable insights on the interplay between SARS-CoV-2 coronavirus disease severity and HLA, including HLA-C*04:01. Even though the WHO has downgraded the COVID-19 emergency status, viral variants of SARS-CoV-2 remain a seasonal concern for individuals with weak and/or compromised immunity and the human population in general. Established HLA associations with disease prognosis, such as the one we are summarizing in this letter, can inform clinical management practice and identify at-risk individuals and populations. This research has been conducted using CGEn's HostSeq Databank (Project ID: DACO-5), funded by the Government of Canada through Genome Canada. The study is also supported by the Canadian Institutes of Health Research (CIHR) [PJT-183608, I.B.]. The authors declare no conflict of interest. The data that support the findings of this study are openly available in Zenodo at https://doi.org/10.5281/zenodo.10463788.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".