The N764K and N856K mutations in SARS-CoV-2 Omicron BA.1 S protein generate potential cleavage sites for SKI-1/S1P protease
Bibliographic record
Abstract
Abstract Spike (S) protein is a key protein in coronaviruses life cycle. SARS-CoV-2 Omicron BA.1 variant of concern (VoC) presents an exceptionally high number of 30 substitutions, 6 deletions and 3 insertions in the S protein. Recent works revealed major changes in the SARS-CoV-2 Omicron biological properties compared to earlier variants of concern (VoCs). Here, these major changes could be explained, at least in part, by the mutations N764K and/or N856K in S2 subunit. These mutations were not previously detected in other VoCs. N764K and N856K generate two potential cleavage sites for SKI-1/S1P serine protease, known to cleave viral envelope glycoproteins. The new sites where SKI-1/S1P could cleave S protein might impede the exposition of the internal fusion peptide for membrane fusion and syncytia formation. Based on the human protein atlas, SKI-1/S1P protease is not found in lung tissues (alveolar cells type I/II and endothelial cells), but present in bronchus and nasopharynx. This may explain why Omicron has change of tissue tropism. Viruses have evolved to use several host proteases for cleavage/activation of envelope glycoproteins. Mutations that allow viruses to change of protease may have a strong impact in host range, cell and tissue tropism, and pathogenesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".