MétaCan
Menu
Back to cohort
Record W4404079324 · doi:10.1093/nar/gkae1030

StreptomeDB 4.0: a comprehensive database of streptomycetes natural products enriched with protein interactions and interactive spectral visualization

2024· article· en· W4404079324 on OpenAlexaff
Yue Feng, Ammar Qaseem, Aurélien F. A. Moumbock, Pascal A Kirchner, Conrad V. Simoben, Yvette I. Malange, Smith B. Babiaka, Mingjie Gao, Stefan Günther

Bibliographic record

VenueNucleic Acids Research · 2024
Typearticle
Languageen
FieldMedicine
TopicMicrobial Natural Products and Biosynthesis
Canadian institutionsUniversity of Toronto
FundersAlbert-Ludwigs-Universität FreiburgChina Scholarship CouncilDeutsche Forschungsgemeinschaft
KeywordsBiologyVisualizationComputational biologyNatural (archaeology)DatabaseBioinformaticsComputer scienceData mining

Abstract

fetched live from OpenAlex

Streptomycetes remain an important bacterial source of natural products (NPs) with significant therapeutic promise, particularly in the fight against antimicrobial resistance. Herein, we present StreptomeDB 4.0, a substantial update of the database that includes expanded content and several new features. Currently, StreptomeDB 4.0 contains over 8500 NPs originating from ∼3900 streptomycetes, manually annotated from ∼7600 PubMed-indexed peer-reviewed articles. The database was enhanced by two in-house developments: (i) automated literature-mined NP-protein relationships (hyperlinked to the CPRiL web server) and (ii) pharmacophore-based NP-protein interactions (predicted with the ePharmaLib dataset). Moreover, genome mining was supplemented through hyperlinks to the widely used antiSMASH database. To facilitate NP structural dereplication, interactive visualization tools were implemented, namely the JSpecView applet and plotly.js charting library for predicted nuclear magnetic resonance and mass spectrometry spectral data, respectively. Furthermore, both the backend database and the frontend web interface were redesigned, and several software packages, including PostgreSQL and Django, were updated to the latest versions. Overall, this comprehensive database serves as a vital resource for researchers seeking to delve into the metabolic intricacies of streptomycetes and discover novel therapeutics, notably antimicrobial agents. StreptomeDB is publicly accessible at https://www.pharmbioinf.uni-freiburg.de/streptomedb.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.016
Threshold uncertainty score0.053

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0030.001
Meta-epidemiology (broad)0.0030.002
Bibliometrics0.0100.010
Science and technology studies0.0010.001
Scholarly communication0.0040.002
Open science0.0020.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0160.016

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.036
GPT teacher head0.363
Teacher spread0.326 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations14
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueNucleic Acids ResearchSame topicMicrobial Natural Products and BiosynthesisFrench-language works237,207