Bibliographic record
Abstract
MS.1 Improving the robustness, throughput and comprehensiveness of quantitative proteomics Michael J. MacCoss(1), Brian Searle(1,2), Lindsay Pino(1), Deanna Plubell(1), Danielle Faivre(1), Gennifer Merrihew(1), Jarrett Egertson(1), Sonia Ting(1), Brendan MacLean(1) (1)University of Washington, (2)Institute for Systems Biology Our goal is to develop high throughput method for sampling peptides with a mass spectrometer that can be used as a quantitative measure of the phenotype. To do this we would like a tandem mass spectrometry (MS/MS) method that can comprehensively sample all peptides in a sample continuously throughout the chromatographic elution. MS/MS acquired using data independent acquisition (DIA) offers significant advantages in terms of selectivity, sensitivity, and dynamic range over a single stage of mass analysis. Quantitative analysis using MS/MS has significant technical advantages over MS1 analysis and we are just now in a situation where we can get high selectivity (<4 m/z) across a majority of the mass range (i.e. 400–1000 m/z) using a rapid duty cycle (<3 sec). That said, this does not mean that there are not substantial challenges to overcome. For example, we need methods to assess whether the peptide measurements are quantitative versus just qualitative. Additionally, global methods like proteomics struggle significantly with signal calibration – making it difficult to compare quantitative measurements between batches, labs, and instrument platforms. Given the prevalence of complex proteoforms we need to think carefully about what the desired outcome is of a quantitative proteomics experiment using bottom-up methodologies. Finally, while most labs feel it is important to measure as many proteins and peptides as possible, the complications associated with doing this is non-trivial – ultimately with an increase in the number of analytes measured increases the multiple testing burden and the number of samples required to have the same statistical power. The talk will present the current state of the art of performing quantitative proteomics using DIA. I will present use cases of what we can do, where we think the limitations are, and what work is being done to improve the methods further. MS.2 Complex-centric proteome profiling by SEC-SWATH-MS Isabell Bludau(1), Moritz Heusel(1), George Rosenberger(1,2), Robin Hafen(1), Max Frank(1,3), Amir Banaei-Esfahani(1), Claudia Martelli(1), Charlotte Nicod(1), Peng Xue(1), Yujia Cai(3), Yansheng Liu(4), Ashok Venkitaraman(5), Vihandha Wickramasinghe(6), Hannes Roest(3), Ben Collins(1), Matthias Gstaiger(1), Ruedi Aebersold(1) (1)Institute of Molecular Systems Biology, ETH Zurich, Switzerland, (2)Columbia University, New York, United States, (3)Donnelly Centre for Cellular and Biomolecular Research, University of Toronto, Canada, (4)Yale University, New Haven, United States, (5)Medical Research Council Cancer Unit, University of Cambridge, United Kingdom, (6)Peter MacCallum Cancer Centre, Melbourne, Australia Proteins are major effectors and regulators of biological processes and can elicit multiple functions depending on their interaction with other proteins. Therefore, it is of central interest in systems biology to determine the interactions and cooperation of proteins as a function of cell state. We have therefore developed an integrated experimental and computational technique for detecting in parallel hundreds of protein complexes, as well as changes in their composition and abundance in a single operation. The method consists of size exclusion chromatography (SEC) to fractionate native protein complexes, SWATH/DIA mass spectrometry to precisely quantify the proteins in each SEC fraction, and the computational framework CCprofiler to detect and quantify protein complexes by error-controlled, complex-centric analysis using prior information from generic protein interaction maps (Heusel & Bludau at al., 2019). Application of our workflow to the HEK293 cell line proteome delineates 462 complexes composed of 2,127 protein subunits, entailing 7673 unique protein-protein interactions. Our analysis further provided insights into novel sub-complexes and assembly intermediates of central regulatory complexes such as the proteasome. We have recently extended this workflow to study rearrangements of protein complex assemblies across different cell states, providing insights into assembly changes that are not captured by full proteome analyses. To increase throughput for such comparative SEC-SWATH-MS analyses, we established a fast protocol based on a 21 minute gradient on the EvoSep One HPLC system that enables the measurement of the ca. 65 SEC fractions of a biological sample per day, while minimizing loss of information. Furthermore, we extended our workflow to take advantage of available peptide-level information in the SEC-SWATH-MS data to investigate proteoform-specific complex integration. We expect our workflow to enable novel insights into the interplay between different protein variants and their impact on protein interactions and functionality on an unprecedented, system wide scale. Moritz Heusel*, Isabell Bludau*, George Rosenberger, Robin Hafen, Max Frank, Amir Banaei-Esfahani, Ben Collins, Matthias Gstaiger, Ruedi Aebersold Complex-centric proteome profiling by SEC-SWATH-MS Molecular Systems Biology Jan 14;15(1):e8438 (2019). MS.3 Aligning label-based discovery and global DIA validation proteomics to explore bacterial virulence phenotypes Stuart J. Cordwell(1), Lok Man(1), Joel A. Cain(1) (1)Charles Perkins Centre and School of Life and Environmental Sciences, The University of Sydney, Aust The use of proteomics to inform subsequent biological validation studies requires substantial rigor in the analytical approach to ensure that the most important leads are followed. Our laboratory explores virulence determinants including an N-linked glycosylation (pgl) system and nutrient transporters in the gastrointestinal pathogen, Campylobacter jejuni. C. jejuni is a Gram negative, spiral and micro-aerophilic bacterium with a sequenced genome containing ∼1620 genes. Target identification is based on the response of the proteome to environmental conditions that mimic the host, including bile salts, low iron, mucin availability and growth temperatures. The proteomics workflow includes parallel label-based liquid chromatography/tandem mass spectrometry (LC-MS/MS; minimum 3 biological growth replicates) using TMT and/or iTRAQ labelling, and system-wide validation using data independent analysis (DIA-SWATH-MS; minimum duplicate additional biological replicates). We routinely quantify ∼80–90% of the predicted C. jejuni NCTC11168 genome using label-based LC-MS/MS (2 peptides; <1% FDR), and ∼65–75% of these can be validated by DIA-SWATH-MS. Here, we will discuss the correlation between large-scale datasets in this biological system and how they facilitate subsequent studies, as well as highlight poorly or non-correlating data. In each case, we show how validated changes in the C. jejuni proteome reflect 'functional reality' that can be determined by molecular genetics, virulence assays and intra- and extra-cellular metabolomics. MS.4 Advanced algorithms to assess and improve quantitative suitability in large DIA datasets Sebastian Vaca(1), Karen Christianson(1), Nicholas Shulman(2), Karsten Krug(1), Brendan X. MacLean(2), Michael J. MacCoss(2), Steven A. Carr(1), Jacob D. Jaffe(1) (1)Broad Institute of MIT and Harvard, Cambridge, MA, (2)University of Washington Genome Sciences, Seattle, WA Data-Independent Acquisition (DIA) is a technique that promises to comprehensively detect and quantify all peptides above an instrument's limit of detection. Several software tools to analyze DIA data have been developed in recent years. However, several challenges still remain, like confidently identifying peptides, defining integration boundaries, dealing with interference for selected transitions, and scoring and filtering of peptide signals in order to control false discovery rates. In practice, a visual inspection of the signals is still required, which is impractical with large datasets. Avant-garde is a new tool to refine DIA (and PRM) by removing interfered transitions, adjusting integration boundaries and scoring peaks to control the FDR. Unlike other tools where MS experiments are scored independently from each other, Avant-garde uses a novel data-driven scoring strategy. DIA signals are refined by learning from the data itself, using all measurements in all samples together to achieve the best optimization. We evaluated the performance of Avant-garde and the results clearly showed that it is capable of improving the selectivity, accuracy, and reproducibility of the quantification results in very complex biological matrices. We have further shown that it can evaluate the suitability of a peak to be used for quantification reaching the same results obtained with manual validation. Avant-garde is envisioned as a tool complementary to existing DIA analysis engines that aims to establish the strongest foundation for subsequent analysis of quantitative MS data. Our workflow was applied to a large cohort of phosphopeptide-enriched samples. The dataset presented here spanned over 6 cell lines, with 90 drug perturbations, employing drugs that span the epigenetic-, neuro- and phosphosignaling-space. The analysis of this data enabled the confident quantification of more than 5000 phosphopeptides in each of the more than 1700 DIA runs. Cumulatively we quantified more than 22000 phosphopeptides in this dataset. Our tool improved the data completeness across the sample set and was robust to retention time shifts. MS.5 Parallel accumulation — serial fragmentation combined with data-independent acquisition (diaPASEF) Florian Meier(1), Andreas-David Brunner(1), Max Frank(2), Annie Ha(2), Eugenia Voytik(1), Stephanie Kaspar-Schoenefeld(3), Markus Lubeck(3), Oliver Raether(3), Ruedi A
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.680 | 0.603 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".