MétaCan
Menu
Back to cohort

Using MIxS: An Implementation Report from Two Metagenomic Information Systems

2017· article· en· W2749169945 on OpenAlexaffabout
Joel L. Sachs, Luke Thompson, Nazir El-Kayssi, Satpal Bilkhu

Bibliographic record

VenueBiodiversity Information Science and Standards · 2017
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsAgriculture and Agri-Food Canada
Fundersnot available
KeywordsMetadataMetagenomicsWorkflowComputer scienceOntologyData scienceHuman Microbiome ProjectInteroperabilitySample (material)Information retrievalData miningWorld Wide WebDatabaseBiology

Abstract

fetched live from OpenAlex

MIxS (Minimum Information about any Sequence) (Yilmaz et al. 2011) is a metadata standard of the Genomics Standards Consortium (GSC), designed to make sequence data findable, accessible, and interoperable. It contains fields for recording physical and chemical characteristics of the sampling environment, geographical and habitat information, and other metadata about the sample and its provenance, which are critical for downstream intepretation of data derived from the sample. We will present our experience implementing MIxS in two metagenomic information systems – the Earth Microbiome Project (EMP) and the Government of Canada (GoC) Ecobiomics project. The EMP (Gilbert et al. 2014) is an ongoing effort to crowdsource environmental microbiome samples from around Earth, then sequence and analyze them using a standardized workflow. The EMP has aggregated and sequenced over 50,000 samples, which are queryable using a publicly available catalogue. A meta-analysis of the first 25,000 samples is currently in review. MIxS and the Environment Ontology (ENVO) (Buttigieg et al. 2016) have been useful in structuring environmental metadata from EMP studies. For the particular application of the EMP meta-analysis, however, several issues were encountered: often there are multiple possible 'correct' assignments to the biome, feature, and material fields; the fields are not hierarchical, limiting logical organization; and the primary ecological factors differentiating microbial communites are not captured. In response to these challenges, the EMP team worked with the ENVO team to devise a new hierarchical structure, the EMP ontology (EMPO), that captures the primary axes along which microbial communities tend to be structured (host-associated or not, saline or not). EMPO is an application ontology, with a formally defined W3C Web Ontology Language (OWL) document mapping to existing ontologies, enabling reuse by the microbial ecology community. Ecobiomics is a joint project of multiple GoC departments and involves the complete workflow, from sampling in a variety of aquatic, soil, and benthic environments, through sample prep, DNA extraction, library prep, sequencing, and analysis. In contrast to the EMP—where some of the samples and metadata had been collected before the establishment of the MIxS standards—the Ecobiomics project has been able to create metadata profiles for each sub-project to conform to, extend, and build, upon the existing MIxS standards. Despite these two different contexts, EMP and Ecobiomics encountered a number of common issues that prevented a complete implementation of MIxS. These issues include ambiguous term names and definitions; inconsistencies amongst the environmental packages; non-standard ways of dealing with units; and a number of issues surrounding ENVO (the Environment Ontology), which is required for filling out the mandatory MIxS fields "Environmental material", "Biome", and "Environmental feature". We will describe these issues, and, more generally, the successes and challenges of our implementations.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.029
metaresearch head score (Gemma)0.039
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Reporting · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.971
Threshold uncertainty score0.153

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0290.039
Meta-epidemiology (narrow)0.0030.004
Meta-epidemiology (broad)0.0010.004
Bibliometrics0.0030.003
Science and technology studies0.0020.001
Scholarly communication0.0100.012
Open science0.0050.016
Research integrity0.0030.006
Insufficient payload (model declined to judge)0.0190.021

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.051
GPT teacher head0.369
Teacher spread0.318 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainReporting
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2017
Admission routes2
Has abstractyes

Explore more

Same venueBiodiversity Information Science and StandardsSame topicBiomedical Text Mining and OntologiesFrench-language works237,207