Notice bibliographique
Résumé
Nextclade Web 2.10.0, Nextclade CLI 2.10.0 (2023-01-24) Add motifs search Nextclade datasets can now be configured to search for motifs in the translated sequences, given a regular expression. At the same time, we released new versions of the following Influenza datasets, which use this feature to detect glycosylation motifs: Influenza A H1N1pdm HA (flu_h1n1pdm_ha), with reference MW626062 Influenza A H3N2 HA (flu_h3n2_ha), with reference EPI1857216 If you run the analysis with the latest version of these datasets, you can find the results in the glycosylaiton column or field of output files or in "Glyc." column in Nextclade Web. If you want to configure your own datasets for motifs search, see an example configuration in the aaMotifs property of virus_properties.json of these datasets: link. Allow to chose columns written into CSV and TSV outputs You can now select a subset of columns to be included into CSV and TSV output files of Nextclade Web (available in the "Download" dialog) and Nextclade CLI (available with --output-csv and --output-tsv). You can either chose individual columns or categories of related columns. In Nextclade Web, in the "Download" dialog, click "Configure columns", then check or uncheck columns or categories you want to keep. Note that this configuration persists across different Nextclade runs. In Nextclade CLI, use --output-columns-selection flag. This flag accepts a comma-separated list of column names and/or column category names. Individual columns and categories can be mixed together. You can find a list of column names in the full output file. The following categories are currently available: all, general, ref-muts, priv-muts, errs-warns, qc, primers, dynamic. Another way to receive both lists is to add a non-existent or misspelled name to the list. The error message will then display all possible columns and categories. Add URL parameter for running analysis of example sequences You can now launch the analysis of example sequences (as provided by the dataset) in Nextclade Web, by using the special keyword example in the input-fasta URL parameter. For example, navigating to this URL will run the analysis of example SARS-CoV-2 sequences (same as choosing "SARS-CoV-2" and then clicking "Load example" in the UI): https://clades.nextstrain.org/?dataset-name=sars-cov-2&input-fasta=example This could useful for example for testing new datasets: https://clades.nextstrain.org/?dataset-url=http://example.com/my-dataset-dir&input-fasta=example Commit history (click to expand) Instructions 📥 Nextclade CLI & Nextalign CLI can be downloaded from the links in the "Assets" section just below. There click "Show all" to show more options. Note the difference between "nextalign" and "nextclade" files. 🌐 Nextclade Web is available at https://clades.nextstrain.org 🐋 Docker images are available at DockerHub 📚 To understand how it all works, make sure to read the Documentation
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,007 |
| Méta-épidémiologie (sens strict) | 0,005 | 0,006 |
| Méta-épidémiologie (sens large) | 0,004 | 0,004 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,006 | 0,006 |
| Science ouverte | 0,008 | 0,005 |
| Intégrité de la recherche | 0,003 | 0,007 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,305 | 0,467 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».