MétaCan
Menu
Retour à la cohorte
Enregistrement W2985622460 · doi:10.1074/mcp.tir119.001752

Assessing Protein Sequence Database Suitability Using De Novo Sequencing

2019· article· en· W2985622460 sur OpenAlexaboutno aff
Richard S. Johnson, Brian C. Searle, Brook L. Nunn, Jason M. Gilmore, Molly Phillips, Chris T. Amemiya, Michelle Heck, Michael J. MacCoss

Notice bibliographique

RevueMolecular & Cellular Proteomics · 2019
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenomics and Phylogenetic Studies
Établissements canadiensnon disponible
Organismes subventionnairesOffice of Extramural Research, National Institutes of HealthNational Institute of General Medical Sciences
Mots-clésComputational biologySequence (biology)Sequence databaseProtein sequencingComputer scienceBiologyDatabasePeptide sequenceGeneticsGene

Résumé

récupéré en direct d'OpenAlex

The analysis of samples from unsequenced and/or understudied species as well as samples where the proteome is derived from multiple organisms poses two key questions. The first is whether the proteomic data obtained from an unusual sample type even contains peptide tandem mass spectra. The second question is whether an appropriate protein sequence database is available for proteomic searches. We describe the use of automated de novo sequencing for evaluating both the quality of a collection of tandem mass spectra and the suitability of a given protein sequence database for searching that data. Applications of this method include the proteome analysis of closely related species, metaproteomics, and proteomics of extinct organisms. The analysis of samples from unsequenced and/or understudied species as well as samples where the proteome is derived from multiple organisms poses two key questions. The first is whether the proteomic data obtained from an unusual sample type even contains peptide tandem mass spectra. The second question is whether an appropriate protein sequence database is available for proteomic searches. We describe the use of automated de novo sequencing for evaluating both the quality of a collection of tandem mass spectra and the suitability of a given protein sequence database for searching that data. Applications of this method include the proteome analysis of closely related species, metaproteomics, and proteomics of extinct organisms. Matching tandem mass spectra to database-derived sequences is routine, and a variety of software pipelines are available (1Eng J.K. Searle B.C. Clauser K.R. Tabb D.L. A face in the crowd: recognizing peptides through database search.Mol. Cell. Proteomics. 2011; 10: 1-9Abstract Full Text Full Text PDF Scopus (126) Google Scholar). However, atypical and new types of samples can be problematic. For example, it can be difficult to evaluate sample preparation methods using LCMS/MS data from fossils, soil, or glacial meltwaters, when such samples are new to the research community, DNA is unattainable, or there is no obvious protein database to search. If the genome of the organism under study is unknown, or if the sample comes from multiple unknown species (i.e. metaproteomics) the suitability of a chosen sequence database can be difficult to evaluate (2Timmins-Schiffman E. May D.H. Mikan M. Riffle M. Frazar C. Harvey H.R. Noble W.S. Nunn B.L. Critical decisions in metaproteomics: Achieving high confidence protein annotations in a sea of unknowns.ISME J. 2017; 11: 309-314Crossref PubMed Scopus (53) Google Scholar). In the case of species whose genomes are unsequenced, data analysis typically employs a protein sequence database from a taxonomically related species, where the hope is that the sequences are mostly identical (3Cilia M. Tamborindeguy C. Rolland M. Howe K. Thannhauser T.W. Gray S. Tangible benefits of the aphid Acyrthosiphon pisum genome sequencing for aphid proteomics: Enhancements in protein identification and data validation for homology-based proteomics.J. Insect Physiol. 2011; 57: 179-190Crossref PubMed Scopus (15) Google Scholar). For metaproteomics, the standard approach is to sequence the DNA in the sample, assemble a metagenome, and translate it into protein sequences as a FASTA formatted file (4Ruggles K.V. Krug K. Wang X. Clauser K.R. Wang J. Payne S.H. Fenyo D. Zhang B. Mani D.R. Methods, tools and current perspectives in proteogenomics.Mol. Cell. Proteomics. 2017; 16: 959-981Abstract Full Text Full Text PDF PubMed Scopus (88) Google Scholar). The hope is that there are no mistakes in assembly and translation to protein sequence, and that the translated metagenome accurately represents the metaproteome under study. Once a FASTA file is created, a database search is performed and the number of identifications is reported, yet it is not clear how many high-quality tandem mass spectra failed to make a match because the peptide sequence was not represented in the FASTA file. If there is more than one sequence database to choose from (e.g. because of different gene assembly methods), the one with the largest number of identifications over a range of false discovery rates is the optimal choice. Although this is a valid way to proceed, one cannot know if a low number of identifications is because of bad tandem mass spectrometry (MS/MS) 1The abbreviations used are:MS/MStandem mass spectrometryCIDcollision induced dissociationABCammonium bicarbonateDDAdata dependent acquisitionPSMpeptide-spectrum matchFDRfalse discovery rateTPtrue positiveFPfalse positiveFNfalse negativeTICtotal ion currentDIAdata independent acquisition. 1The abbreviations used are:MS/MStandem mass spectrometryCIDcollision induced dissociationABCammonium bicarbonateDDAdata dependent acquisitionPSMpeptide-spectrum matchFDRfalse discovery rateTPtrue positiveFPfalse positiveFNfalse negativeTICtotal ion currentDIAdata independent acquisition. data, an inappropriate or insufficient sequence database, or both. Here we propose and evaluate a simple solution that uses automated de novo sequencing to evaluate both tandem mass spectrum quality and sequence database suitability. tandem mass spectrometry collision induced dissociation ammonium bicarbonate data dependent acquisition peptide-spectrum match false discovery rate true positive false positive false negative total ion current data independent acquisition. tandem mass spectrometry collision induced dissociation ammonium bicarbonate data dependent acquisition peptide-spectrum match false discovery rate true positive false positive false negative total ion current data independent acquisition. De novo sequencing is the concept of deriving a peptide sequence from a tandem mass spectrum without use of a sequence database (5Ma B. Johnson R. De novo sequencing and homology searching.Mol. Cell. Proteomics. 2012; 11: 1-16Abstract Full Text Full Text PDF Scopus (129) Google Scholar). Before the existence of sequence databases (or ready access to powerful computers), de novo sequencing was the only approach to interpret tandem mass spectra of peptides. The degree to which one can successfully derive a sequence is dependent on spectral quality. Specifically, spectra that can be sequenced possess a contiguous series of sequencing ions of the same type (e.g. b- and y-type ions), which turns out to be the case for low energy collision induced dissociation (CID) of peptides. In the automated de novo sequencing deriving sequences in a the of the sequences and with the use of high mass such as or The automated de novo sequencing B. peptide de novo sequencing PubMed Scopus Google the in that the software can de novo sequences than it to the data in than a database search of the same The question we with de novo Here we describe an approach that to tandem mass spectra quality and the suitability of FASTA for database searches. many high de novo sequences in a data that many high-quality tandem mass spectra of peptides are Johnson and uses of automated de novo peptide sequencing tandem mass PubMed Scopus Google Scholar). of de novo sequencing with database search from the same data file and with the same database search can be used to the suitability of a chosen FASTA file. The is to de novo sequences to the FASTA file under a standard database search of the FASTA and the of the de novo sequences with the FASTA file a number of to the FASTA file with the de novo that the FASTA file is for use in a database search. Here we that automated de novo sequencing using a simple and of evaluating both tandem mass spectra and FASTA file quality. A was obtained from A C. was as C. B. Noble W.S. of proteomics for the and of C. gene PubMed Scopus Google Scholar). of performed as Johnson J. M. the and of the 10: Scopus Google Scholar). of E. of C. de M. M. E. E. R. C. B. D. J. the genome and of the extinct in the of PubMed Scopus Google M. D. J. S. sequencing of PubMed Scopus Google J. M. B. C. S. M. genome sequence of a from DNA PubMed Scopus Google was in of in ammonium bicarbonate and for on of the to in The was and the to total protein using a A and a and in with and The was from and was out the in The was from the in to the of where For both the of was on the of of this with of and the sample from was the DNA is in in the and of with for an to and for using of C. was and with and as In the the samples to using on a and in of mass spectrometry was performed on a or mass to of sample was from the a with to a of a rate of and using a total of of the was with a with the same that was in a and in with a the using a of in over over a rate of The mass using with the using data dependent acquisition in the or The was and for tandem mass spectrometry (MS/MS) the ion or the was a of the spectra using a of and collision energy of or collision energy of was for using file with B. R. D. D.L. S. B. B. J. K. D. B. C. D. K. C. B. J. S. J. B. K. J. Tabb D.L. A for mass spectrometry and 2012; PubMed Scopus Google Scholar). De novo sequences using the B. peptide de novo sequencing PubMed Scopus Google and the used and no the appropriate was used as a a file that the de novo sequence and a sequence of the was performed high confidence peptide sequence search with the de novo The database search used J.K. an tandem mass spectrometry sequence database search 2012; E. R. to the of peptide identifications and database PubMed Scopus Google using the D. E. B. B. J.K. B. R. A of the 10: PubMed Scopus Google Scholar). was used to de novo sequences for spectra that in a database search peptide spectrum with false discovery rates and than If the de novo sequence for of the high confidence database sequence, it was as this on this de novo sequences with of or into a protein sequence that was to the appropriate FASTA file For the analysis of a proteome FASTA used from a number of mostly species B. C. S. and X. For the analysis of a C. proteome FASTA used from a species C. C. D. and The in the was the number of protein In FASTA from and C. the and the data and FASTA from the Noble D.H. E. Mikan Harvey H.R. E. Nunn B.L. Noble W.S. for of samples using PubMed Scopus Google Scholar). The metagenome, and FASTA and data was a proteome FASTA file from and a FASTA file from data was a proteome from C. and data was C. as well as a of sequence data for in and from was two derived from gene of the D. genome and where the contains sequences of and and DNA from glacial was used to make metagenome FASTA are in the to of FASTA was a of database performed with using FASTA that high de novo In was for to in to a de novo sequence or to of multiple de novo if such than sequence, the be The was to and the in was a of or as appropriate the proteome was using the for of of the protein and of the protein The search was The file from was using a to with the that sequences derived from a FASTA file can a than a de novo when the two sequences are a sequence the de novo sequence to for one or two in the first and second to be The the first and second sequences used to this The for this approach is that we sequences are of and that the first a because of is when a de novo sequence is than a FASTA file sequence for the of peptide on to the peptide The suitability of this is in which the match The the and through a of the is used in to a that can be used as a than that of the of the when a de novo sequence than a FASTA Once the file E. R. to the of peptide identifications and database PubMed Scopus Google was used to the where the was in with the of only the of high confidence sequences for as to for evaluating high sequences and to a FASTA and to analysis and are available The de novo sequencing an sequence for sequence that and where the number a how to interpret this peptides in with of and high low mass on a mass using For peptides using a database search of a FASTA file. The and both to be and with the sequences from the same spectra. simple de novo and database sequences is not De novo sequences are not are obvious be because of the to and a of low or sequencing ions in of sequence in an A automated de novo sequencing be to of sequence and a sequence with a high In bad spectra insufficient ions to of the sequence a low tandem mass spectra (e.g. from not in high de novo to de novo sequences is it a to to database because a be a way to make using mass Johnson database de novo peptide sequencing tandem mass 11: PubMed Scopus Google B.C. S. M. D. identification of and sequence using a for de novo sequencing PubMed Scopus (129) Google Scholar). In this example, the two are which is a because of the of ions the first and second the of and be as because of the of is when an ion in in the de novo sequence when it is the of a where the de novo sequences with with the database spectra are in In there was in the of the yet of the be the two of the be the in in this the de novo sequence and The database sequence contains a of the de novo sequence this is no in this was a case where two peptides database search out one of the peptides and sequenced the In a search of the de novo sequence an match to a where of the However, for the if a that of the peptide mass is the de novo sequence is with database search with low false discovery rates which to be was from a of a on a mass that in the ion or and high or low or ion are the number of de novo sequences or a given sequence are the number of de novo sequences or a given are the number of de novo sequences a given be that a de novo sequence only to be to mass of the peptide the number of de novo sequences whose sequence the to in a and the for the same data. is if ion are with high and The of type on the mass is for high mass and is for low the types of data used to and not to the of the and number of de novo sequences obtained for and with mass data for a different For example, if a of is a sequence of was to over of the de novo sequences when using In this there de novo sequences out of a total of spectra in a database of which a sequence in the range of to to be an appropriate for de novo sequences for for data. data was using with high mass mass of the can be difficult to evaluate and a new sample preparation method or a new type of sample, when there is no obvious FASTA file to search. typically the of the total ion current the and the number of tandem mass spectra using the of it to evaluate the of de novo sequences database where can of the spectra are high quality peptide the of and the number of de novo sequences that quality of ions in a file from a using a The in such a as to the appropriate mass or an to when only of the ions there was a in the number of high de novo a of different of that and to the of of high-quality peptide spectra derived from samples that difficult to and not in protein this one that there are many peptide spectra in the data, it to be how many are because of an of the which to the of spectra that can be to peptide sequences in the database to the that are de novo de novo sequences are into a protein with a protein sequence is to the of the FASTA file to be the FASTA file is using a standard database search such as using the and are to both FASTA database sequences and de novo For a given one the number of de novo and database sequences to the database quality. this database one to be to the sequence in a and to when a database derived sequence and a de novo are (e.g. with the to the is the database sequence to the de novo to be a make this we that the two are For spectra that two the sequence are with the of the to and from high to low (i.e. from the two to of the where a sequence is second a de novo sequence, one use the as the A more approach a of the total (e.g. where the be obtained the (e.g. of the way A was that a on and sequences that are this if are second to a de novo this high confidence from a of a and of to a search of the FASTA file to which was the de novo sequences for high confidence the of the it is in this that sequences are and that a de novo sequence a that this is because of and with The novo is to the peptide and the same the two are and in are obtained using a different data obtained from a of C. The in employs to rates and using this file. The number of de novo sequences that than database sequence using the derived from the two can be for as of the of the de novo In this the number of de novo sequences can be with the total FASTA and de novo sequences and the database quality is as the the number of de novo sequences is the database quality is the number of de novo sequences is the database quality is we this method to to evaluate be high or low can where proteomic data are for a species that a protein sequence database (or In one search a FASTA file from a closely related The question as to whether the chosen database is The in was from a variety of related and organisms. The data in and when proteomic data is FASTA file protein sequence a database of which to a FASTA The be because of sequence and the database was the de novo of the sequences to the FASTA this analysis for closely related and in of the sequences to FASTA this to for the more related from of the sequences match to and the for more related and even to and to as FASTA from more related are well the obtained from the B. a and not proteomic data was the approach that of a FASTA file. and the obtained when proteomic data from C. is FASTA from of the In this an and C. FASTA file with de novo sequences in and of the sequences to the FASTA file. A of and that is not a to a related FASTA file. The closely related C. is in the same yet only of the sequences be to this FASTA file. The genome sequence two taxonomically related is to Although FASTA from the taxonomically related organism is the in on the of sequence of this are using proteomic data obtained from from the extinct a from a and a from a the was the species related to for which a proteome was In this the search was performed with and as in to as was to be a protein in the a database search was performed on the FASTA file to which was the de novo peptide of the sequences to the FASTA sequences that in a for many of we that or the The search was that of to and to as was as a than to the search given the of of in a The number of sequences to the FASTA file from to that of the over The total number of sequences from to which is because of the to match to when is a that of the be and out of sequences in the the proteome for and using as a the of FASTA file derived sequences to high be because are and are protein of novo analysis of FASTA protein sequence of data from and The and the number of and sequences that to de novo than the FASTA file sequences of as of as in a new For the from the the for which a proteome FASTA file is available is from the In this of the sequences FASTA sequences A was performed on a different In this two FASTA the FASTA and a file of sequences from The number of sequences to the two FASTA and The data that the available FASTA are not for standard database the analysis of samples derived from a of unknown and evaluating a protein sequence database to search is (2Timmins-Schiffman E. May D.H. Mikan M. Riffle M. Frazar C. Harvey H.R. Noble W.S. Nunn B.L. Critical decisions in metaproteomics: Achieving high confidence protein annotations in a sea of unknowns.ISME J. 2017; 11: 309-314Crossref PubMed Scopus (53) Google Scholar). typically are databases only of from or species (e.g. are not to be in one can DNA assemble the into a metagenome, and translate the metagenome sequences to a metaproteome (4Ruggles K.V. Krug K. Wang X. Clauser K.R. Wang J. Payne S.H. Fenyo D. Zhang B. Mani D.R. Methods, tools and current perspectives in proteogenomics.Mol. Cell. Proteomics. 2017; 16: 959-981Abstract Full Text Full Text PDF PubMed Scopus (88) Google Scholar). The assembly of a metagenome from DNA sequence and translation to metaproteome can May D.H. E. Mikan Harvey H.R. E. Nunn B.L. Noble W.S. for of samples using PubMed Scopus Google the concept of a In this peptide of a from a simple translation of the DNA in the May the number of identifications from two different and the In this search from the analysis of FASTA a database from (i.e. protein sequences translated from metagenome from DNA sequencing of sample, and the databases derived from the same DNA sequencing data. The database is an assembly of translated sequences obtained from sequencing only of which FASTA and using de novo sequence analysis which that the database in for both the and The metagenome databases to be to the where the FASTA file is to the database or the metagenome, the May D.H. E. Mikan Harvey H.R. E. Nunn B.L. Noble W.S. for of samples using PubMed Scopus Google Scholar). on organisms are for of a for In to the the analysis of in is a that a FASTA file protein sequences from multiple organisms. a file was from and and of this FASTA file that of the sequences be to this FASTA file. FASTA file was to include there was only a in the of sequences to the new FASTA file FASTA be which is because the of the genome is S. K. S. M. C. D. K. K. S. R. J. M. D. D. K. X. Johnson S. B.L. S. C. M. M. D. of the of a 2017; Scopus Google Scholar). However, the current database is for Johnson J. M. the and of the 10: Scopus Google Johnson R. S. X. J. M. the in the of the 2017; PubMed Scopus Google Johnson R. M. of and in of the of PubMed Scopus Google S. S. Johnson R. K. and to the and the in the 2017; Scopus Google Scholar). A more under study is the analysis of and in the (i.e. and the or (i.e. with and The same samples to DNA sequencing to derive metagenome databases for the of protein sequence databases for mass the of de novo analysis of FASTA using the data. In tandem mass spectra de novo than from translated that samples it that many of the peptides be from However, FASTA sequences yet the with the FASTA file novo analysis of FASTA from obtained from from a samples obtained from and from the or sequences in a new Although a software can automated de novo is available as and more software is and and de novo sequencing no is with data obtained using mass of whether ion or For low mass data, to using be because of the from the of both b- and y-type ions typically is that the low data that used for used than the available in is on de novo sequencing of peptides derived from In well on peptides where the ions are using high and for low using For proteome of the data quality typically that the is and that many spectra for is more difficult to make an of data from unusual samples glacial where the not be and one is not that the ions are derived from peptides. A for an of data be an of the number of spectra for which a high de novo sequence can be The be to that peptide spectra are not because of and which is a database search a FASTA file to which sequences of A of is the and suitability of the FASTA file to which are not be Here we how to use de novo sequencing to evaluate the suitability of a chosen FASTA file de novo sequences to a FASTA a database search on this and of the match to the FASTA sequences with the de novo spectra match to the sequences of the can match when such FASTA sequences a species under study a FASTA one typically to use a taxonomically related For example, peptides to over of the C. to C. the not and a for the suitability of a FASTA file is Although this a it not on how to a Matching a low of peptide spectra to a given FASTA file can for a of the of peptides or the of the sequences in the FASTA file. data not match well to the database the database search was performed with (i.e. and one for a match the data be from or For example, if the peptides high de novo sequences that sequence to the mass because of not match to a In the case of a proteome sample that in a for of a that the be However, sample preparation the for low FASTA is that the FASTA file is the The method a way to evaluate FASTA For the data where there FASTA to metagenome, and was that the was the choice. the FASTA file was than the FASTA file for the data, and the FASTA file sequences was than the FASTA file that only sequences for In the sequences to the FASTA file for the data only a the of high-quality genome annotations for the of proteomic is the question of to with the where the de novo sequences are than are to sequences or peptides is to homology-based (e.g. using the de novo given that de novo sequences are not such homology be to sequencing into Johnson database de novo peptide sequencing tandem mass 11: PubMed Scopus Google and B.C. S. M. D. identification of and sequence using a for de novo sequencing PubMed Scopus (129) Google two of the where the approach is to match of de novo sequences in the of However, two are or given the of data a this a that be is given the and of A second an search to to sequences or peptides D. and peptide identification in mass 2017; PubMed Scopus Google and a be to a search using an FASTA database J. E. Clauser K. J.K. M. FASTA PubMed Scopus Google Scholar). this approach for evaluating FASTA can be used on data independent acquisition data. on the spectra the D. B. M. for acquisition PubMed Scopus Google and in the same Searle B. De novo of the on and Scholar). The mass spectrometry proteomics data to the M. S. S. Wang M. The in the in proteomics data 2017; PubMed Scopus Google the J. M. S. J. M. E. J. J. S. J. E. M. The database and related tools and in for PubMed Scopus Google with the and We for in using We E. for the with

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,029
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,034
Tête enseignante GPT0,279
Écart entre enseignants0,245 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations43
Publié2019
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueMolecular & Cellular ProteomicsMême sujetGenomics and Phylogenetic StudiesTravaux en français237 207