MétaCan
Menu
← Back to cohort
Record W4415885267 · doi:10.1038/s41598-025-22523-z

Manually weighted taxonomy classifiers improve species-specific rumen microbiome analysis compared to unweighted or average weighted taxonomy classifiers

2025· article· en· W4415885267 on OpenAlexaff
Ryukseok Kang, Zhongtang Yu, Hanbeen Kim, Jakyeom Seo, Minseok Kim, Tansol Park

Bibliographic record

VenueScientific Reports · 2025
Typearticle
Languageen
FieldAgricultural and Biological Sciences
TopicRuminant Nutrition and Digestive Physiology
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsRefSeqWeightingMicrobiomeTaxonomic rankBiological classificationClassifier (UML)Taxonomy (biology)MetagenomicsRandom forest

Abstract

fetched live from OpenAlex

Previous research has demonstrated that applying taxonomic weights to shotgun metagenomic data can improve species identification in 16S rRNA gene-based microbiome analysis. However, such an approach does not allow for accurate analysis of samples collected from less studied habitats, such as rumen. In the present study, we developed a method to incorporate taxonomic weights based on relative abundance of species identified from shotgun sequencing and amplicon sequencing data derived from rumen. Using this weighting method, we evaluated latest versions of five prominent databases-SILVA, Greengenes2 (GG2), RDP, NCBI RefSeq, and GTDB-against the BLAST 16S rRNA database, assessing classification counts, fully classified ratios (proportion of ASVs classified to a known genus and species), and error rates. Our results indicated that providing taxonomic weights partially increased classification counts and fully classified ratios, although the extent of improvement varied across databases. A reduction in error rates was also observed compared to the unweighted taxonomy classifier (P < 0.05). While GG2 and SILVA struggled with accurate classification at the species level owing to their inherent database characteristics, GTDB consistently improved all metrics using the manually weighted taxonomy classifier, achieving up to an 8% error rate reduction at the species level. NCBI RefSeq and RDP also exhibited remarkable improvement in the classification counts and fully classified ratios, along with error rate reductions by up to 47% at the species level. These findings demonstrate that amplicon sequencing datasets can enhance rumen microbiome analyses through effective weighting methods. While SILVA is commonly used in metataxonomic analyses of the rumen microbiome, we recommend NCBI RefSeq for species-level classification due to its superior accuracy and minimal ambiguous classification (e.g., "uncultured" or "sp.") in future metataxonomic studies.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.015
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.006
Threshold uncertainty score0.029

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0060.015
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0040.002
Science and technology studies0.0010.000
Scholarly communication0.0020.003
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.037
GPT teacher head0.237
Teacher spread0.200 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueScientific Reports→Same topicRuminant Nutrition and Digestive Physiology→French-language works237,207→