MétaCan
Menu
Back to cohort
Record W2952844973 · doi:10.1074/mcp.tir118.001132

DIAlignR Provides Precise Retention Time Alignment Across Distant Runs in DIA and Targeted Proteomics

2019· article· en· W2952844973 on OpenAlexaff
Shubham Gupta, Sara Ahadi, Wenyu Zhou, Hannes Röst

Bibliographic record

VenueMolecular & Cellular Proteomics · 2019
Typearticle
Languageen
FieldChemistry
TopicAdvanced Proteomics Techniques and Applications
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsRobustness (evolution)Computer sciencePairwise comparisonElutionPattern recognition (psychology)Artificial intelligenceChemistryChromatography

Abstract

fetched live from OpenAlex

Sequential Windowed Acquisition of All Theoretical Fragment Ion Mass Spectra (SWATH-MS) is widely used for proteomics analysis given its high throughput and reproducibility, but ensuring consistent quantification of analytes across large-scale studies of heterogeneous samples such as human plasma remains challenging. Heterogeneity in large-scale studies can be caused by large time intervals between data acquisition, acquisition by different operators or instruments, and intermittent repair or replacement of parts, such as the liquid chromatography column, all of which affect retention time (RT) reproducibility and, successively, performance of SWATH-MS data analysis. Here, we present a novel algorithm for RT alignment of SWATH-MS data based on direct alignment of raw MS2 chromatograms using a hybrid dynamic programming approach. The algorithm does not impose a chronological order of elution and allows for alignment of elution-order-swapped peaks. Furthermore, allowing RT mapping in a certain window around a coarse global fit makes it robust against noise. On a manually validated dataset, this strategy outperformed the current state-of-the-art approaches. In addition, on real-world clinical data, our approach outperformed global alignment methods by mapping 98% of peaks compared with 67% cumulatively. DIAlignR reduced alignment error up to 30-fold for extremely distant runs. The robustness of technical parameters used in this pairwise alignment strategy is also demonstrated. The source code is released under the BSD license at https://github.com/Roestlab/DIAlignR. Sequential Windowed Acquisition of All Theoretical Fragment Ion Mass Spectra (SWATH-MS) is widely used for proteomics analysis given its high throughput and reproducibility, but ensuring consistent quantification of analytes across large-scale studies of heterogeneous samples such as human plasma remains challenging. Heterogeneity in large-scale studies can be caused by large time intervals between data acquisition, acquisition by different operators or instruments, and intermittent repair or replacement of parts, such as the liquid chromatography column, all of which affect retention time (RT) reproducibility and, successively, performance of SWATH-MS data analysis. Here, we present a novel algorithm for RT alignment of SWATH-MS data based on direct alignment of raw MS2 chromatograms using a hybrid dynamic programming approach. The algorithm does not impose a chronological order of elution and allows for alignment of elution-order-swapped peaks. Furthermore, allowing RT mapping in a certain window around a coarse global fit makes it robust against noise. On a manually validated dataset, this strategy outperformed the current state-of-the-art approaches. In addition, on real-world clinical data, our approach outperformed global alignment methods by mapping 98% of peaks compared with 67% cumulatively. DIAlignR reduced alignment error up to 30-fold for extremely distant runs. The robustness of technical parameters used in this pairwise alignment strategy is also demonstrated. The source code is released under the BSD license at https://github.com/Roestlab/DIAlignR. In translational research, protein biomarkers and therapeutic targets are usually discovered by data-driven methods such as by linking protein abundance patterns with disease conditions. A large sample cohort is essential in these studies as substantial biological variability exists in the population and enough statistical power is required to identify disease-specific events (1Uzozie A.C. Aebersold R. Advancing translational research and precision medicine with targeted proteomics.J. Proteomics. 2018; 189: 1-10Crossref PubMed Scopus (53) Google Scholar, 2Surinova S. Schiess R. Hüttenhain R. Cerciello F. Wollscheid B. Aebersold R. On the development of plasma protein biomarkers.J. Proteome Res. 2011; 10: 5-16Crossref PubMed Scopus (247) Google Scholar). Blood plasma is a good source of clinical information of a patient as it can be obtained noninvasively and proteins from affected tissue can potentially leak into the blood. Plasma samples, unfortunately, are highly challenging for proteomic analysis due to the diversity of peptides within the samples and high dynamic range of plasma proteins (3Nigjeh E.N. Chen R. Brand R.E. Petersen G.M. Chari S.T. von Haller P.D. Eng J.K. Feng Z. Yan Q. Brentnall T.A. Pan S. Quantitative proteomics based on optimized data-independent acquisition in plasma analysis.J. Proteome Res. 2017; 16: 665-676Crossref PubMed Scopus (28) Google Scholar). Therefore, quantification of plasma proteins requires a highly reproducible reduction of complexity and measurement within a wide dynamic range. The situation is exacerbated across large-scale studies, which makes development of plasma biomarker challenging (2Surinova S. Schiess R. Hüttenhain R. Cerciello F. Wollscheid B. Aebersold R. On the development of plasma protein biomarkers.J. Proteome Res. 2011; 10: 5-16Crossref PubMed Scopus (247) Google Scholar, 3Nigjeh E.N. Chen R. Brand R.E. Petersen G.M. Chari S.T. von Haller P.D. Eng J.K. Feng Z. Yan Q. Brentnall T.A. Pan S. Quantitative proteomics based on optimized data-independent acquisition in plasma analysis.J. Proteome Res. 2017; 16: 665-676Crossref PubMed Scopus (28) Google Scholar). In the past two decades, mass spectrometry (MS)-based proteomics has made rapid advances with a high degree of innovation in obtaining identification and quantification of proteins in various biological samples (2Surinova S. Schiess R. Hüttenhain R. Cerciello F. Wollscheid B. Aebersold R. On the development of plasma protein biomarkers.J. Proteome Res. 2011; 10: 5-16Crossref PubMed Scopus (247) Google Scholar, 4Schubert O.T. Röst H.L. Collins B.C. Rosenberger G. Aebersold R. Quantitative proteomics: Challenges and opportunities in basic and applied research.Nat. Protoc. 2017; 12: 1289-1294Crossref PubMed Scopus (139) Google Scholar). Targeted proteomics methods, specifically selected reaction monitoring (SRM), can provide high reproducibility across multiple runs. However, it is limited by low throughput and can measure abundance of only a few tens to low hundreds of proteins per study (1Uzozie A.C. Aebersold R. Advancing translational research and precision medicine with targeted proteomics.J. Proteomics. 2018; 189: 1-10Crossref PubMed Scopus (53) Google Scholar, 5Röst H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar). we approach for targeted analysis of data-independent acquisition used under the acquisition of all mass retention of identification used under the acquisition of all mass retention of identification data, which can of peptides in large-scale clinical studies H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, Navarro P. S. Röst L. R. Aebersold R. Targeted data of the by data-independent A for consistent and Proteomics. PubMed Scopus Google Scholar). this in the clinical provide of samples across various conditions. has reproducible quantification of proteins in a biomarker study on and (1Uzozie A.C. Aebersold R. Advancing translational research and precision medicine with targeted proteomics.J. Proteomics. 2018; 189: 1-10Crossref PubMed Scopus (53) Google Scholar, P. Gillet Röst H.L. Rosenberger G. Collins B.C. S. M. Aebersold R. mass of tissue samples into PubMed Scopus Google and has the to a of samples a large of monitoring of a patient (1Uzozie A.C. Aebersold R. Advancing translational research and precision medicine with targeted proteomics.J. Proteomics. 2018; 189: 1-10Crossref PubMed Scopus (53) Google Scholar). data-independent acquisition under the liquid chromatography error retention time chromatograms acquisition of all mass retention time of identification data-independent acquisition under the liquid chromatography error retention time chromatograms acquisition of all mass retention time of identification In in are selected for a range and MS2 of of all selected The data can be by using a approach H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, H.L. Rosenberger G. Navarro P. Gillet L. O.T. Collins B.C. Malmström Malmström L. Aebersold R. targeted analysis of data-independent acquisition PubMed Scopus Google or a approach B. M. for data-independent acquisition proteomics.Nat. Methods. 12: PubMed Scopus Google Scholar). to be of and protein quantification in samples H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, P. Gillet B. Röst H.L. L. Rosenberger G. Y. Aebersold R. S. A study for 2016; PubMed Scopus Google Scholar, Y. Collins B.C. Gillet G. Aebersold R. Quantitative variability of plasma proteins in a human PubMed Scopus Google Scholar). obtaining reproducible and robust analysis of clinical plasma samples is challenging with as large in the of proteins in are H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, P. Gillet B. Röst H.L. L. Rosenberger G. Y. Aebersold R. S. A study for 2016; PubMed Scopus Google Scholar, Y. Collins B.C. Gillet G. Aebersold R. Quantitative variability of plasma proteins in a human PubMed Scopus Google Scholar). of the variability is the retention time (RT) between and plasma elution In by and of the peptides RT of between technical the robustness of quantification (3Nigjeh E.N. Chen R. Brand R.E. Petersen G.M. Chari S.T. von Haller P.D. Eng J.K. Feng Z. Yan Q. Brentnall T.A. Pan S. Quantitative proteomics based on optimized data-independent acquisition in plasma analysis.J. Proteome Res. 2017; 16: 665-676Crossref PubMed Scopus (28) Google Scholar). also in and identification of the peptides (3Nigjeh E.N. Chen R. Brand R.E. Petersen G.M. Chari S.T. von Haller P.D. Eng J.K. Feng Z. Yan Q. Brentnall T.A. Pan S. Quantitative proteomics based on optimized data-independent acquisition in plasma analysis.J. Proteome Res. 2017; 16: 665-676Crossref PubMed Scopus (28) Google Scholar). data analysis peptides to a RT L. B. R. F. a retention time for targeted measurement of 12: PubMed Scopus Google or R. L. in the targeted analysis of data-independent acquisition and its on identification and 2016; 16: PubMed Scopus Google with to a this chromatograms from MS2 are obtained for usually multiple in which makes analysis challenging. elution time can be for H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, R. alignment in and A 16: PubMed Scopus Google Scholar). A in RT is as a which is using a between two R. alignment in and A 16: PubMed Scopus Google Scholar). However, this not be and specifically distant to a are the elution order of two peptides is across two R. alignment in and A 16: PubMed Scopus Google Scholar, M. retention time with of the in PubMed Scopus Google Scholar, L. S. A hybrid retention time alignment algorithm for SWATH-MS 2016; 16: PubMed Scopus Google Scholar). is in studies and in large-scale clinical studies in which data acquisition a of are methods in the for in RT alignment in and proteomics the development of SWATH-MS R. alignment in and A 16: PubMed Scopus Google and, on chromatograms of and multiple for data analysis using Scopus Google Scholar, R. G. alignment by and dynamic programming as a for of liquid spectrometry PubMed Scopus Google Scholar, S.T. alignment of time Y. L. in information Scholar, A for time alignment of PubMed Scopus Google Scholar, P. M. Aebersold R. B. for mass Proteomics. PubMed Scopus Google Scholar, retention time alignment for spectrometry PubMed Scopus Google Scholar, F. R. alignment based on selected mass for Proteome Res. PubMed Scopus Google A dynamic programming approach for the alignment of peaks in multiple spectrometry PubMed Scopus Google Scholar, R. M. M. M. A for analysis of PubMed Scopus (139) Google Scholar, alignment for multiple liquid spectrometry PubMed Scopus Google Scholar, M. S. F. An alignment algorithm for Proteomics. 12: PubMed Scopus Google or a of alignment of proteomics data by PubMed Scopus Google Scholar, M. M. P. and retention time alignment for multiple spectrometry 13: PubMed Scopus Google Scholar). usually a global pairwise alignment using dynamic programming on raw chromatograms A for time alignment of PubMed Scopus Google Scholar, P. M. Aebersold R. B. for mass Proteomics. PubMed Scopus Google Scholar, retention time alignment for spectrometry PubMed Scopus Google Scholar, F. R. alignment based on selected mass for Proteome Res. PubMed Scopus Google Scholar, alignment of proteomics data by PubMed Scopus Google or on A dynamic programming approach for the alignment of peaks in multiple spectrometry PubMed Scopus Google Scholar, R. M. M. M. A for analysis of PubMed Scopus (139) Google which methods using samples, alignment of proteomics data by PubMed Scopus Google Scholar, M. M. P. and retention time alignment for multiple spectrometry 13: PubMed Scopus Google used to RT alignment However, of these on data, and the alignment are by all In SWATH-MS MS2 data high and are reproducible across multiple runs. research on RT alignment of has on MS2 MS2 using L. S. A hybrid retention time alignment algorithm for SWATH-MS 2016; 16: PubMed Scopus Google or the to a global by in S. or by a approach H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, B.C. quantification for data acquisition mass spectrometry using 2018; provide in of high or Furthermore, the global not for as a RT between two peptides L. S. A hybrid retention time alignment algorithm for SWATH-MS 2016; 16: PubMed Scopus Google Scholar). we present RT alignment algorithm these of algorithm does not and is of the raw MS2 from targeted proteomics approach dynamic programming to mapping between chromatograms information as peaks around the RT alignment of the alignment of elution-order-swapped peaks. is also of using a global alignment for it robust against noise. DIAlignR can of between of global and provide to our source code and our at https://github.com/Roestlab/DIAlignR. our on a manually validated of chromatograms and performance also it on selected plasma selected from a heterogeneous cohort across our algorithm outperformed global alignment methods and of by distant it also which is not using global alignment methods H.L. Rosenberger G. Navarro P. Gillet L. O.T. Collins B.C. Malmström Malmström L. Aebersold R. targeted analysis of data-independent acquisition PubMed Scopus Google Scholar, L. B. R. F. a retention time for targeted measurement of 12: PubMed Scopus Google Scholar). are for mapping which are to runs. can be used for as the is we used a H.L. Rosenberger G. Navarro P. Gillet L. O.T. Collins B.C. Malmström Malmström L. Aebersold R. targeted analysis of data-independent acquisition PubMed Scopus Google of from In these selected manually H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar). of for two and from the peaks of the and, from Therefore, for performance of the DIAlignR against global alignment approaches. of the are in the peaks selected from the dataset, this peaks with SWATH-MS on human plasma samples in from to peptides of plasma samples on a used with using a to a from with A with with of plasma on a analysis using Acquisition on a with a and window acquisition methods in with the of pairwise two from and are in the peaks not only peaks with a low for performance Therefore, peaks of with a and required to be present in all runs. chromatograms of the selected and using H.L. Rosenberger G. Navarro P. Gillet L. O.T. Collins B.C. Malmström Malmström L. Aebersold R. targeted analysis of data-independent acquisition PubMed Scopus Google Scholar, G. M. Liu Y. L. Röst H.L. S. Collins B.C. Aebersold R. of and protein error in large-scale targeted data-independent acquisition Methods. 2017; PubMed Scopus Google and B. R. S. L. B. B. B. T.A. L. B. F. P. S. B. M. P. A for mass spectrometry and proteomics.Nat. PubMed Scopus Google with The RT of the peptides in all is in the The chromatograms are on A of is as of and human plasma plasma from acquisition of of selected for of used for and of selected per for of in a In targeted proteomics or SWATH-MS is using or In we using at for analysis O.T. Röst H.L. Collins B.C. Rosenberger G. Aebersold R. Quantitative proteomics: Challenges and opportunities in basic and applied research.Nat. Protoc. 2017; 12: 1289-1294Crossref PubMed Scopus (139) Google Scholar). or A of or chromatograms is a which to the given a using for the of which the raw data for our alignment A can be a of The of the between from and can be a has and has and time in and as in the between all time can be as a The is as a measure and can be selected by the In our we in for chromatograms such as and H.L. Rosenberger G. Navarro P. Gillet L. O.T. Collins B.C. Malmström Malmström L. Aebersold R. targeted analysis of data-independent acquisition PubMed Scopus Google Scholar, alignment of proteomics data by PubMed Scopus Google Scholar). the between all and data information and between two time elution from the data of the is by a in in the of the two be as in with the from all of is as and the of in and A of is in However, to the of a as used for A in the using dynamic which to RT alignment from to and dynamic programming a of the in the is by alignment to and can to a the alignment is highly from a global or the alignment robust against and in order to information from the global we in our algorithm to the using global alignment such as error of the fit to a of in the and of it with a to alignment within a time window to global and large The alignment by all from the of the to the of it using dynamic programming R. G. of proteins and Google Scholar). and not mapping as these be chromatograms around the elution by retention time for Therefore, alignment of a global alignment of MS2 approach and chromatograms for or a of is a as it the Therefore, with for a of In this which a for of R. G. of proteins and Google Scholar). The alignment using is in The time of such alignment is A data-driven approach is to from the the alignment to the time chromatograms as in of MS2 chromatograms of has a time of order chromatograms of different can be Therefore, we for different peptides to a for are various parameters used in A of these parameters is in a for and used the of peaks within and RT alignment error as used the manually validated H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google to DIAlignR to the current state-of-the-art of H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google which a of high peaks to a or alignment RT from to as as for optimized from used in and of the in S. Scholar). to the between two H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar, in S. Scholar). to a global fit mapping are in the and RT error by against the of the S. H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google and the of the of peptides within a certain RT used as a measure of of the alignment are not for the human plasma the with low of used for Here, we present algorithm for alignment only raw MS2 data from targeted proteomics or for RT the performance of our we the of parameters on the of the alignment of from the H.L. Liu Y. D'Agostino G. Zanella M. Navarro P. Rosenberger G. Collins B.C. Gillet L. Testa G. Malmström L. Aebersold R. TRIC: An automated alignment strategy for reproducible protein quantification in targeted proteomics.Nat. Methods. 2016; 13: PubMed Scopus Google Scholar). we the performance for different of The with as a measure the of peptides for all RT error of this measure of peaks with the under the of all and the of the used in dynamic In the is as a of the of the of not a on the of peaks within a certain RT the to of peaks within The of for RT the selected as the for the a of a of the of the of on the alignment and as the of alignment with the alignment from using all compared with for only However, the from to two with for using all a not for this but the two algorithm the using a global alignment the alignment in a certain window by around the global fit from the alignment with a peaks compared with with the An of such alignment is in and in which the two high the alignment the in the alignment the the alignment as in the a dataset, we compared DIAlignR to current alignment In of of peptides and alignment alignment by DIAlignR outperformed and methods and On the S. dataset, DIAlignR error by compared with the state-of-the-art alignment only of all peaks within of the RT compared with for with parameters of all and of peptides from manually validated S. and from heterogeneous human plasma the plasma data, peaks with are used for of the of peaks within of peaks within of peaks within two of peaks per within human plasma in a the of on the performance of the alignment compared with the dataset, the and human plasma to S. alignment methods, we performance for compared with However, the performance of the the performance of DIAlignR robustness to sample for the of alignment across multiple we the of peaks as the alignment within for the with low for the alignment compared with the consistent in its performance In of the precision of the alignment global alignment methods as the a under the for for of we a RT with the which the alignment with the and and on the dataset, DIAlignR in of of alignment and of peaks across a range of different RT and runs. we in the global between the two methods to the alignment error for pairwise alignment of and alignment outperformed in of all in of performance in the On DIAlignR reduced the RT error by with a of our peaks compared with by optimized within However, in we on the dataset, methods with which be due to the low complexity of a sample and the high of the data as within two on the performance on the S. dataset, we the performance of our algorithm on the data from large-scale SWATH-MS on human a challenging as the data a of with intermittent repair of the and replacement of the with a from of the selected at and peptides used for our we not manually used as the our alignment algorithm with the on a highly heterogeneous human plasma dataset, we our approach of peaks compared with by with a error of as in All pairwise performance using alignment we in the performance of our on the alignment of on the two different for on different alignment with of peaks compared with for by the DIAlignR performance for highly heterogeneous the approach we not only DIAlignR outperformed on but also the performance for DIAlignR for compared with the performance of alignment we to its across the of the of peaks in all DIAlignR of peaks on within peaks on a substantial alignment error using which be reduced by the performance on we the alignment error for pairwise alignment for The of alignment error for compared with for DIAlignR the precision of RT alignment with our approach. of we DIAlignR outperformed in of all and in in of these peaks to be due to by of the alignment approach on the heterogeneous human plasma validated its consistent and RT alignment In liquid RT is from to However, we this be for different peptides and in of retention order M. retention time with of the in PubMed Scopus Google Scholar). In such a two peptides are in order in elution order in our approach not of order of elution and we DIAlignR be of of we elution order from the heterogeneous and distant plasma runs. the alignment for such by alignment we specifically at the alignment of the as it has the of of on from on The from peptides for this is in of the peptides around the global fit of on the from of the of peptides elution order of peptides of peptides in at of of the is in In in the elution order to be peptides RT in from only RT between two caused the peptides to in a different The be with a global alignment which in the be by our alignment the peaks from to the of peptides for alignment 98% peaks compared with which to only DIAlignR of the error by up to which not by by the and to be the of by due to the and for RT and between has a in and it has of as proteomics large-scale analysis of human However, on data, and few are can the information present in MS2 by targeted methods or In this we a novel algorithm used the raw chromatograms to RT alignment for targeted proteomics and algorithm to peaks across a of and alignment compared to current state-of-the-art the algorithm and a hybrid which used a global alignment to the to in hybrid approach the of with a allows the to on global or on our not alignment on raw The dynamic programming approach is essential for obtaining a alignment as distant also for peaks. to the alignment made our algorithm and the robustness of global alignment on a dataset, DIAlignR outperformed a global alignment or the current state-of-the-art alignment a two of and of by using all which also to O.T. Röst H.L. Collins B.C. Rosenberger G. Aebersold R. Quantitative proteomics: Challenges and opportunities in basic and applied research.Nat. Protoc. 2017; 12: 1289-1294Crossref PubMed Scopus (139) Google Scholar). the DIAlignR approach on data, we also of the global alignment alignment we and the DIAlignR algorithm error from to we our to in or sample global alignment to the novel alignment be to in sample and in large-scale our algorithm on a large-scale SWATH-MS of human plasma On this dataset, DIAlignR reduced RT alignment error from to which is a current state-of-the-art approach outperformed methods and the of peaks within of acquisition time or repair between two DIAlignR RT alignment which has the to identification and quantification manually of a as in also in the of a of our to RT of it as our hybrid approach also global alignment can be and be used to peaks. this can be to chromatograms by and The in RT alignment with alignment of with a DIAlignR on pairwise alignment of selected peptides per The in the alignment of the of the alignment using dynamic However, this with the of peptides in the and be on a our approach is for large-scale heterogeneous targeted proteomics studies are by different and data are or a mapping in such a challenging the of elution order of alignment chronological order of elution and, we substantial with the global and reduction of error using hybrid approach these peptides as it on of to peaks. is peptides this is the is O.T. Röst H.L. Collins B.C. Rosenberger G. Aebersold R. Quantitative proteomics: Challenges and opportunities in basic and applied research.Nat. Protoc. 2017; 12: 1289-1294Crossref PubMed Scopus (139) Google and in such our global alignment RT alignment has multiple in the of proteomics for large-scale identification and of a large of analytes are two of as at to on RT present a can of data by between analytes across large of samples, for and also this can be by proteomics to identification and to the chromatograms and by are on under code are to for data acquisition and to the heterogeneous plasma also for on alignment using dynamic with

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.045
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.005
GPT teacher head0.225
Teacher spread0.220 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations36
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueMolecular & Cellular ProteomicsSame topicAdvanced Proteomics Techniques and ApplicationsFrench-language works237,207