MétaCan
Menu
Back to cohort
Record W2158439765 · doi:10.1186/1741-7007-5-44

Broadening the horizon – level 2.5 of the HUPO-PSI format for molecular interactions

2007· article· en· W2158439765 on OpenAlexafffund
Samuel Kerrien, Sandra Orchard, Luisa Montecchi‐Palazzi, Bruno Aranda, A. F. Quinn, Nisha Vinod, Gary D. Bader, Ioannis Xénarios, Jérôme Wojcik, David James Sherman, Mike Tyers, John J Salama, Susan Moore, Arnaud Céol, Andrew Chatr‐aryamontri, Matthias Oesterheld, Volker Stümpflen, Łukasz Salwiński, Jason Nerothin, Ethan Cerami, Michael E. Cusick, Marc Vidal, Michael K. Gilson, J. T. Armstrong, Peter Woollard, Christopher W.V. Hogue, David Eisenberg, Gianni Cesareni, Rolf Apweiler, Henning Hermjakob

Bibliographic record

VenueBMC Biology · 2007
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsMount Sinai HospitalUniversity of TorontoLunenfeld-Tanenbaum Research Institute
FundersNational Institutes of HealthEuropean CommissionGenome CanadaOntario GenomicsNational Institute of General Medical SciencesOntario Genomics InstituteU.S. Department of Energy
KeywordsCollaboratoryHuman proteome projectComputer scienceSchema (genetic algorithms)XMLData scienceWorld Wide WebComputational biologyBiologyInformation retrievalProteomics

Abstract

fetched live from OpenAlex

BACKGROUND: Molecular interaction Information is a key resource in modern biomedical research. Publicly available data have previously been provided in a broad array of diverse formats, making access to this very difficult. The publication and wide implementation of the Human Proteome Organisation Proteomics Standards Initiative Molecular Interactions (HUPO PSI-MI) format in 2004 was a major step towards the establishment of a single, unified format by which molecular interactions should be presented, but focused purely on protein-protein interactions. RESULTS: The HUPO-PSI has further developed the PSI-MI XML schema to enable the description of interactions between a wider range of molecular types, for example nucleic acids, chemical entities, and molecular complexes. Extensive details about each supported molecular interaction can now be captured, including the biological role of each molecule within that interaction, detailed description of interacting domains, and the kinetic parameters of the interaction. The format is supported by data management and analysis tools and has been adopted by major interaction data providers. Additionally, a simpler, tab-delimited format MITAB2.5 has been developed for the benefit of users who require only minimal information in an easy to access configuration. CONCLUSION: The PSI-MI XML2.5 and MITAB2.5 formats have been jointly developed by interaction data producers and providers from both the academic and commercial sector, and are already widely implemented and well supported by an active development community. PSI-MI XML2.5 enables the description of highly detailed molecular interaction data and facilitates data exchange between databases and users without loss of information. MITAB2.5 is a simpler format appropriate for fast Perl parsing or loading into Microsoft Excel.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.019
metaresearch head score (Gemma)0.037
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.028
Threshold uncertainty score0.100

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0190.037
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0050.005
Science and technology studies0.0020.002
Scholarly communication0.0110.016
Open science0.0050.010
Research integrity0.0040.008
Insufficient payload (model declined to judge)0.0280.031

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.089
GPT teacher head0.356
Teacher spread0.268 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations276
Published2007
Admission routes2
Has abstractyes

Explore more

Same venueBMC BiologySame topicBiomedical Text Mining and OntologiesFrench-language works237,207