In silico prediction of COVID-19 test efficiency with DinoKnot
Bibliographic record
Abstract
Abstract The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is a novel coronavirus spreading across the world causing the disease COVID-19. The diagnosis of COVID-19 is done by quantitative reverse-transcription polymer chain reaction (qRT-PCR) testing which utilizes different primer-probe sets depending on the assay used. Using in silico analysis we aimed to determine how the secondary structure of the SARS-CoV-2 RNA genome affects the interaction between the reverse primer during qRT-PCR and how it relates to the experimental primer-probe test efficiencies. We introduce the program DinoKnot (Duplex Interaction of Nucleic acids with pseudoKnots) that follows the hierarchical folding hypothesis to predict the secondary structure of two interacting nucleic acid strands (DNA/RNA) of similar or different type. DinoKnot is the first program that utilizes stable stems in both strands as a guide to find the structure of their interaction. Using DinoKnot we predicted the interaction of the reverse primers used in four common COVID-19 qRT-PCR tests with the SARS-CoV-2 RNA genome. In addition, we predicted how 12 mutations in the primer/probe binding region may affect the primer/probe ability and subsequent SARS-CoV-2 detection. While we found all reverse primers are capable of interacting with their target area, we identified partial mismatching between the SARS-CoV-2 genome and some reverse primers. We predicted three mutations that may prevent primer binding, reducing the ability for SARS-CoV-2 detection. We believe our contributions can aid in the design of a more sensitive SARS-CoV-2 test. Author summary The current testing for the disease COVID-19 that is caused by the novel cornonavirus SARS-CoV-2 uses oligonucleotides called primers that bind to specific target regions on the SARS-CoV-2 genome to detect the virus. Our goal was to use computational tools to predict how the structure of the SARS-CoV-2 RNA genome affects the ability of the primers to bind to their target region. We introduce the program DinoKnot (Duplex interaction of nucleic acids with pseudoknots) that is able to predict the interactions between two DNA or RNA molecules. We used DinoKnot to predict the efficiency of four common COVID-19 tests, and the effect of mutations in the SARS-CoV-2 virus on ability of the COVID-19 tests in detecting those strains. We predict partial mismatching between some primers and the SARS-CoV-2 genome but that all primers are capable of interacting with their target areas. We also predict three mutations that prevent primer binding and thus SARS-CoV-2 detection. We discuss the limitations of the current COVID-19 testing and suggest the design of a more sensitive COVID-19 test that can be aided by our findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".