Tracking Co-Occurrence of N501Y, P681R, and Other Key Mutations in SARS-CoV-2 Spike for Surveillance
Bibliographic record
Abstract
The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has produced five variants of concern (VOC) to date. The important spike mutation ‘N501Y’ is common to Alpha, Beta, Gamma, and Omicron VOC, while the ‘P681R’ is key to Delta’s spread. We have analysed circa 10 million SARS-CoV-2 genome sequences from the world’s largest repository, ‘Global Initiative on Sharing All Influenza Data (GISAID)’, and demonstrated that these two mutations have co-occurred on the spike ‘D614G’ mutation background at least 5767 times from 12 May 2020 to 28 April 2022. In contrast, the Y501-H681 combination, which is common to Alpha and Omicron VOC, is present in circa 1.1 million entries. Over half of the 5767 co-occurrences were in France, Turkey, or US (East Coast), and the rest across 88 other countries; 36.1%, 3.9%, and 4.1% of the co-occurrences were Alpha’s Q.4, Gamma’s P.1.8, and Omicron’s BA.1.1 sub-lineages acquiring the P681R; 4.6% and 3.0% were Delta’s AY.5.7 sub-lineage and B.1.617.2 lineage acquiring the N501Y; the remaining 8.2% were in other variants. Despite the selective advantages individually conferred by N501Y and P681R, the Y501-R681 combination counterintuitively did not outcompete other variants in every instance we have examined. While this is a relief to worldwide public health efforts, in vitro and in vivo studies are urgently required in the absence of a strong in silico explanation for this phenomenon. This study demonstrates a pipeline to analyse combinations of key mutations from public domain information in a systematic manner and provide early warnings of spread. The study here demonstrates the usage of the pipeline using the key mutations N501Y, P681R, and D614G of SARS-CoV-2.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".