Initial Guidelines for Manuscripts Employing Data-independent Acquisition Mass Spectrometry for Proteomic Analysis
Bibliographic record
Abstract
Proteomic research began largely as an approach for characterizing sample compositions, but most contemporary studies involve a quantitative aspect. Quantification enables comparing sample classes (e.g. healthy versus disease) to uncover markers of dysregulation, or comparing protein pull-down experiments to mock pull-downs to determine specific interaction partners. For small-scale comparisons, isotopic labeling, whether introduced metabolically or chemically, is very effective and allows comparison of multiple samples mixed together. However, for comparing a larger number of samples (a dozen or more), label-free strategies are often the most practical option. Reproducible and accurate quantification of a large number of protein and peptide analytes across a large panel of samples remains a singular goal of the proteomics field in general. Data-independent acquisition mass spectrometry (DIA-MS) is a set of strategies that aim to provide comprehensive coverage and quantification of components in complex peptide mixtures. DIA-MS was developed to circumvent the issues of irreproducible selection of analytes for fragmentation analysis associated with data-dependent acquisition (DDA) and limited analyte coverage (typically m/z range, most commonly broken down into a series of isolated wide m/z range windows (1Purvine S. Eppel J.T. Yi E.C. Goodlett D.R. Shotgun collision-induced dissociation of peptides using a time of flight mass analyzer.Proteomics. 2003; 3: 847-850Crossref PubMed Scopus (130) Google Scholar, 2Venable J.D. Dong M.Q. Wohlschlegel J. Dillin A. Yates J.R. Automated approach for quantitative analysis of complex peptide mixtures from tandem mass spectra.Nat. Methods. 2004; 1: 39-45Crossref PubMed Scopus (509) Google Scholar, 3Chapman J.D. Goodlett D.R. Masselon C.D. Multiplexed and data-independent tandem mass spectrometry for global proteome profiling.Mass Spectrom. Rev. 2014; 33: 452-470Crossref PubMed Scopus (184) Google Scholar, 4Egertson J.D. Kuehn A. Merrihew G.E. Bateman N.W. MacLean B.X. Ting Y.S. Canterbury J.D. Marsh D.M. Kellmann M. Zabrouskov V. Wu C.C. MacCoss M.J. Multiplexed MS/MS for improved data-independent acquisition.Nat. Methods. 2013; 10: 744-746Crossref PubMed Scopus (207) Google Scholar, 5Moseley M.A. Hughes C.J. Juvvadi P.R. Soderblom E.J. Lennon S. Perkins S.R. Thompson J.W. Steinbach W.J. Geromanos S.J. Wildgoose J. Langridge J.I. Richardson K. Vissers J.P.C. Scanning quadrupole data-independent acquisition, Part A: Qualitative and quantitative characterization.J. Proteome Res. 2018; 17: 770-779Crossref PubMed Scopus (43) Google Scholar). It has seen considerable growth in the last couple of years as instrumentation that can produce high mass accuracy fragmentation spectra at rates in excess of 10 Hz has become widely available. In parallel to development of acquisition methodologies, new analysis software has also emerged to interpret the resulting data. Molecular and Cellular Proteomics has led the proteomics field in establishing rules for minimum information needed to be provided in submitted manuscripts to evaluate results from different analysis strategies, producing guidelines for authors performing data-dependent MSMS analysis (6Bradshaw R.A. Burlingame A.L. Carr S. Aebersold R. Reporting protein identification data: The next generation of guidelines.Mol. Cell. Proteomics. 2006; 5: 787-788Abstract Full Text Full Text PDF PubMed Scopus (203) Google Scholar), targeted proteomics (7Abbatiello S. Ackermann B.L. Borchers C. Bradshaw R.A. Carr S.A. Chalkley R. Choi M. Deutsch E. Domon B. Hoofnagle A.N. Keshishian H. Kuhn E. Liebler D.C. MacCoss M. MacLean B. Mani D.R. Neubert H. Smith D. Vitek O. Zimmerman L. New guidelines for publication of manuscripts describing development and application of targeted mass spectrometry measurements of peptides and proteins.Mol. Cell. Proteomics. 2017; 16: 327-328Abstract Full Text Full Text PDF PubMed Scopus (32) Google Scholar), glycomics/glycoproteomics (8Wells L. Hart G.W. Glycomics: Building upon proteomics to advance glycosciences.Mol. Cell. Proteomics. 2013; 12: 833-835Abstract Full Text Full Text PDF PubMed Scopus (23) Google Scholar), and clinical proteomic studies (9Celis J.E. Carr S.A. Bradshaw R.A. New guidelines for clinical proteomics manuscripts.Mol. Cell. Proteomics. 2008; 7: 2071-2072Abstract Full Text Full Text PDF Scopus (9) Google Scholar). These guidelines have in general elevated the standard of published results. Of late, the journal has published several DIA-MS studies, and it has become evident that even though DIA-MS strategies are still rapidly evolving, a first set of guidelines is required to advise authors on information that should be included in such manuscripts. Hence, in June 2018, the journal organized a meeting of leading researchers in the DIA-MS field in San Diego, CA, to formulate a mutually agreeable set of rules to cover current and anticipated analysis strategies. Representatives from key DIA method, software, and instrument development groups ensured broad community participation. The full list of attendees is provided at the bottom. The guidelines produced from this meeting were opened to a two-month period of public comment, and the final version is now published (http://www.mcponline.org/page/DIA-guidelines) along with this issue of the journal. A companion checklist has also been constructed to assist authors in meeting these guidelines. The journal intends to start implementing these guidelines for relevant manuscripts on March 1. As DIA-MS methods are still developing, it is anticipated that these guidelines will need to evolve over time to encompass new approaches, but having a first set of guidelines in place will provide a framework for ensuring that results published using these approaches are accountable. We would like to thank Steve Carr and Saddiq Zahari for assisting in the organization of the meeting and MCP, Thermo, Waters, and Bruker for providing financial support. The attendees at the meeting were: Chris Adams, BrukerNuno Bandeira, UCSDIsabell Bludau, ETH ZürichAndreas Brunner, Max Planck Institute of BiochemistryAl Burlingame, UCSFSteven Carr, Broad Institute (Co-organizer)Robert Chalkley, UCSF (Co-organizer)Meena Choi, Northeastern UniversityMike Hoopmann, Institute for Systems BiologyJake Jaffe, Broad InstituteBrendan MacLean, University of WashingtonMike MacCoss, University of WashingtonAlexey Nesvizhskii, University of MichiganLukas Reiter, BiognosysHannes Röst, University of TorontoBirgit Schilling, Buck InstituteBrian Searle, Proteome SoftwareStephen Tate, SCIEXStefan Tenzer, Johannes Gutenberg University MainzHans Vissers, Waters CorporationOlga Vitek, Northeastern UniversityJuan Antonio Vizcaino, EMBL-EBISue Weintraub, UT Health San AntonioYue Xuan, Thermo Fischer ScientificSaddiq Zahari, ASBMB
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".