Automated Library Construction and Analysis for High-Throughput Nanopore Sequencing of SARS-CoV-2
Bibliographic record
Abstract
BACKGROUND: To support the implementation of high-throughput pipelines suitable for SARS-CoV-2 sequencing and analysis in a clinical laboratory, we developed an automated sample preparation and analysis workflow. METHODS: We used the established ARTIC protocol with approximately 400 bp amplicons sequenced on Oxford Nanopore's MinION. Sequences were analyzed using Nextclade, assigning both a clade and quality score to each sample. RESULTS: A total of 2179 samples on twenty-five 96-well plates were sequenced. Plates of purified RNA were processed within 12 h, sequencing required up to 24 h, and analysis of each pooled plate required 1 h. The use of samples with known threshold cycle (Ct) values enabled normalization, acted as a quality control check, and revealed a strong correlation between sample Ct values and successful analysis, with 85% of samples with Ct < 30 achieving a "good" Nextclade score. Less abundant samples responded to enrichment with the fraction of Ct > 30 samples achieving a "good" classification rising by 60% after addition of a post-ARTIC PCR normalization. Serial dilutions of 3 variant of concern samples, diluted from approximately Ct = 16 to approximately Ct = 50, demonstrated successful sequencing to Ct = 37. The sample set contained a median of 24 mutations per sample and a total of 1281 unique mutations with reduced sequence read coverage noted in some regions of some samples. A total of 10 separate strains were observed in the sample set, including 3 variants of concern prevalent in British Columbia in the spring of 2021. CONCLUSIONS: We demonstrated a robust automated sequencing pipeline that takes advantage of input Ct values to improve reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".