MétaCan
Menu
Back to cohort
Record W2135976131 · doi:10.1074/mcp.o111.015446

Recommendations for Mass Spectrometry Data Quality Metrics for Open Access Data (Corollary to the Amsterdam Principles)

2011· article· en· W2135976131 on OpenAlexaff
Christopher R. Kinsinger, James Apffel, Mark S. Baker, Xiaopeng Bian, Christoph H. Borchers, Ralph Bradshaw, Mi‐Youn Brusniak, Daniel W. Chan, Eric W. Deutsch, Bruno Domon, Jeff Gorman, Rudolf Grimm, William S. Hancock, Henning Hermjakob, David M. Horn, Christie L. Hunter, Patrik Kolar, Hans‐Joachim Kraus, Hanno Langen, Rune Linding, Robert L. Moritz, Gilbert S. Omenn, Ron Orlando, Akhilesh Pandey, Peipei Ping, Amir Rahbar, Robert Rivers, Sean L. Seymour, Richard J. Simpson, Douglas J. Slotta, Richard Smith, Stephen E. Stein, David L. Tabb, Danilo A. Tagle, John R. Yates, Henry Rodriguez

Bibliographic record

VenueMolecular & Cellular Proteomics · 2011
Typearticle
Languageen
FieldChemistry
TopicAdvanced Proteomics Techniques and Applications
Canadian institutionsUniversity of Victoria
FundersNational Institute of General Medical SciencesNational Human Genome Research Institute
KeywordsComputer scienceData scienceData qualityQuality (philosophy)Data sharingService (business)World Wide WebMedicineBusiness

Abstract

fetched live from OpenAlex

Policies supporting the rapid and open sharing of proteomic data are being implemented by the leading journals in the field. The proteomics community is taking steps to ensure that data are made publicly accessible and are of high quality, a challenging task that requires the development and deployment of methods for measuring and documenting data quality metrics. On September 18, 2010, the United States National Cancer Institute convened the "International Workshop on Proteomic Data Quality Metrics" in Sydney, Australia, to identify and address issues facing the development and use of such methods for open access proteomics data. The stakeholders at the workshop enumerated the key principles underlying a framework for data quality assessment in mass spectrometry data that will meet the needs of the research community, journals, funding agencies, and data repositories. Attendees discussed and agreed up on two primary needs for the wide use of quality metrics: 1) an evolving list of comprehensive quality metrics and 2) standards accompanied by software analytics. Attendees stressed the importance of increased education and training programs to promote reliable protocols in proteomics. This workshop report explores the historic precedents, key discussions, and necessary next steps to enhance the quality of open access data. By agreement, this article is published simultaneously in the Journal of Proteome Research, Molecular and Cellular Proteomics, Proteomics, and Proteomics Clinical Applications as a public service to the research community. The peer review process was a coordinated effort conducted by a panel of referees selected by the journals.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.204
metaresearch head score (Gemma)0.344
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesMetaresearch
DomainCandidate signal: Reporting · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.984
Threshold uncertainty score0.981

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2040.344
Meta-epidemiology (narrow)0.0040.003
Meta-epidemiology (broad)0.0030.006
Bibliometrics0.0140.015
Science and technology studies0.0040.007
Scholarly communication0.0220.022
Open science0.0160.011
Research integrity0.0280.024
Insufficient payload (model declined to judge)0.0180.017

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.334
GPT teacher head0.431
Teacher spread0.097 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainReporting
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations103
Published2011
Admission routes1
Has abstractyes

Explore more

Same venueMolecular & Cellular ProteomicsSame topicAdvanced Proteomics Techniques and ApplicationsFrench-language works237,207