Bibliographic record
Abstract
Nextclade Web 2.10.0, Nextclade CLI 2.10.0 (2023-01-24) Add motifs search Nextclade datasets can now be configured to search for motifs in the translated sequences, given a regular expression. At the same time, we released new versions of the following Influenza datasets, which use this feature to detect glycosylation motifs: Influenza A H1N1pdm HA (flu_h1n1pdm_ha), with reference MW626062 Influenza A H3N2 HA (flu_h3n2_ha), with reference EPI1857216 If you run the analysis with the latest version of these datasets, you can find the results in the glycosylaiton column or field of output files or in "Glyc." column in Nextclade Web. If you want to configure your own datasets for motifs search, see an example configuration in the aaMotifs property of virus_properties.json of these datasets: link. Allow to chose columns written into CSV and TSV outputs You can now select a subset of columns to be included into CSV and TSV output files of Nextclade Web (available in the "Download" dialog) and Nextclade CLI (available with --output-csv and --output-tsv). You can either chose individual columns or categories of related columns. In Nextclade Web, in the "Download" dialog, click "Configure columns", then check or uncheck columns or categories you want to keep. Note that this configuration persists across different Nextclade runs. In Nextclade CLI, use --output-columns-selection flag. This flag accepts a comma-separated list of column names and/or column category names. Individual columns and categories can be mixed together. You can find a list of column names in the full output file. The following categories are currently available: all, general, ref-muts, priv-muts, errs-warns, qc, primers, dynamic. Another way to receive both lists is to add a non-existent or misspelled name to the list. The error message will then display all possible columns and categories. Add URL parameter for running analysis of example sequences You can now launch the analysis of example sequences (as provided by the dataset) in Nextclade Web, by using the special keyword example in the input-fasta URL parameter. For example, navigating to this URL will run the analysis of example SARS-CoV-2 sequences (same as choosing "SARS-CoV-2" and then clicking "Load example" in the UI): https://clades.nextstrain.org/?dataset-name=sars-cov-2&input-fasta=example This could useful for example for testing new datasets: https://clades.nextstrain.org/?dataset-url=http://example.com/my-dataset-dir&input-fasta=example Commit history (click to expand) Instructions 📥 Nextclade CLI & Nextalign CLI can be downloaded from the links in the "Assets" section just below. There click "Show all" to show more options. Note the difference between "nextalign" and "nextclade" files. 🌐 Nextclade Web is available at https://clades.nextstrain.org 🐋 Docker images are available at DockerHub 📚 To understand how it all works, make sure to read the Documentation
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.005 | 0.006 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.008 | 0.005 |
| Research integrity | 0.003 | 0.007 |
| Insufficient payload (model declined to judge) | 0.305 | 0.467 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".