Proteogenomics: Opportunities and Caveats
Notice bibliographique
Résumé
Proteogenomics is a rapidly evolving field at the intersection of genomics, transcriptomics, and proteomics. Whole genome, exome, and RNA sequencing are well-established techniques that can provide information at the DNA and RNA level with excellent sequencing coverage and depth. Although tens of thousands of clinical samples have been sequenced thus far, data integration and interpretation still remain largely incomplete. Recent advances in proteomic technologies have enabled the accurate and almost complete characterization of the proteomes of many tissues and biological fluids. Integration of multiomics data for the accurate annotation and reciprocal refinement of genomic and proteomic models is essentially the goal of proteogenomics. This integrative approach has the potential to provide solid evidence for the translation of previously unknown transcripts. Those transcripts and the respective encoded proteins might be implicated in physiological or pathophysiological processes. Novel reported peptides can represent single amino acid variants, splice variants, gene fusions, RNA editing events, novel open reading frames, translated noncoding RNAs, and pseudogenes, among many others. Proteogenomic platforms can now be used to investigate which of these novel “events” gets translated at the protein level, thereby implicating them as candidate new druggable targets or as new diagnostic or prognostic biomarkers for a wide spectrum of diseases. The potential for such identifications is maximized when both sequencing and raw proteomic data originate from the very same sample under investigation. It is becoming clear that this “sample-specific” approach, and the use of matched customized search databases, is associated with lower false-positive and false-negative identification rates. However, like all areas of active research, proteogenomics in its current state is not free of drawbacks. Major limitations in the field are the sensitivity of the mass spectrometers, the increased false discovery rate for the novel peptide hits, and the inherent biophysical properties that render some peptides undetectable. In this Q&A we discuss with 4 experts in the field the current status of proteogenomics and conditions that have to be met to deliver its promises. What are the key technologies that enabled the development of proteogenomics? Alexey Nesvizhskii: In most cases, especially when studying human or model organisms, proteogenomics is critically dependent on the knowledge (and often aims to refine that knowledge) assembled by large genome and proteome annotation teams that build genome-centric resources such as Ensemble and RefSeq and protein-level resources such as the UniProt knowledge base (UniProtKB).9 Thus, proteogenomics is critically dependent on those efforts and the technologies that they use. With respect to experiment-specific data used as part of proteogenomics studies, on the genomics side it commonly involves next generation sequencing (NGS) data such as exome sequencing and transcriptomics (RNA-Seq). Because proteogenomics is most dependent on the availability of high-quality proteomic data, the technologies that enable sensitive, large-scale proteome profiling are of the highest importance. Tandem mass spectrometry (MS/MS) is the most dominant technology for high-throughput quantitative proteome profiling. The most commonly used strategy is to digest proteins into peptides using an enzyme such as trypsin, followed by MS/MS sequencing of the resulting peptides. This variant of proteomics is called shotgun (or bottom-up) proteomics. Intact proteins can also be analyzed using mass spectrometry (top-down proteomics), and such data can naturally be used as part of proteogenomics studies as well, but for technical reasons top-down proteomics has yet to enter the proteomics mainstream. Ribosomal proofing is a promising technology that provides complementary information at the level of translational products. This technology is at present technically challenging for most laboratories. However, such data are extremely valuable for proteogenomics applications as they allow more direct linkage between transcriptomics and proteomics data. Last but not least, bioinformatics is of critical importance. This includes databases and public repositories for storing raw data, tools for processing the data coming from each technology, and tools for integration and visualization of such data. In particular, the development of proteomics data repositories such as PeptideAtlas and ProteomeXchange was critical for proteogenomics because they provided computational scientists interested in proteogenomics access to large proteomics data sets needed for their work. Thomas Kislinger: Proteogenomics was first described over a decade ago and has played a crucial role in the development and refinement of genome models for a variety of organisms. Renewed interest in this type of an approach is certainly a result of the rapid improvements in sequencing technologies [NGS, exome-Seq, whole genome sequencing (WGS), RNA-Seq], but likely more closely tied to technological innovations allowing for the characterization of deep proteome profiles from small biological/clinical samples. These more comprehensive data sets have renewed hope that, alongside appropriate bioinformatics and statistical frameworks, novel biologically relevant results can be gleaned from proteogenomic data. Jacob Jaffe: The now routine nature of genome sequencing is an important enabler for proteogenomics today. But even when genome sequencing was relatively more difficult, the data could be leveraged for early proteogenomic applications. Today, proteogenomics and gene expression profiling are becoming even more aligned with the development of RNA-Seq and ribosomal footprinting technologies, enabling better mapping to gene products in specific cellular contexts. Andrei Drabovich: Genome-wide next-generation DNA and RNA sequencing, high-resolution mass spectrometry, sample-specific customized protein databases, bioinformatic algorithms, and software tools to integrate multiple omics data sets enable proteogenomics and facilitate its use for practical applications. As an example, high-resolution mass spectrometry with hybrid quadrupole-time-of-flight or quadrupole-Orbitrap instruments allows for high-throughput analysis of many thousand proteins and provides accurate peptide sequencing data, thus reducing the number of false-positive matches. Such methods eventually help in the identification of rare peptide variants, such as cancer-specific missense mutations. What is the additional information that proteogenomics can offer compared to either genomic or proteomic platforms? Alexey Nesvizhskii: Using a narrower definition of the term proteogenomics, proteogenomics studies aim to validate or refine gene models and annotations produced using genomic data. Existing databases such as Ensemble already contain many annotated transcripts whose protein coding potential is unknown, i.e., sequences predicted based on their genome annotation pipelines but with little or no previous experimental evidence of their expression at the protein level. Similarly, the UniProtKB database—the most commonly used reference database for proteomics studies, contains sequences divided into categories (evidence levels), including the most dubious proteins that have been assigned to P4 and P5 categories. By matching tandem mass spectra against the sequences of those proteins, one can identify their presence in biological samples. Thus, proteomics-based evidence in the form of identified peptides, and in combination with other relevant data that may be available (sequence conservation, RNA-Seq, ribosomal profiling, etc.), can be used to improve the annotation of the corresponding transcripts/proteins. As the most desirable scenario, proteogenomics analysis would result in a promotion of a previously questionable transcript to the status of experimentally confirmed protein-coding sequence. Similarly, proteogenomics can provide protein-level evidence for a particular variant or isoform, e.g., an amino acid variant or an alternative splice form. In a broader sense, proteogenomics refers to all sorts of applications involving joint analysis of sample-specific genomics/transcriptomics and proteomics data. Studies in which NGS data, such as exome and especially RNA-Seq transcriptomic data, and proteomics data generated in parallel are becoming increasingly common. In such studies, sample-specific genomic and transcriptomic data can be used to reconstruct the transcriptomes of the samples under investigation. Such transcriptomes are inherently noisy, i.e., they contain sequence reads suggesting thousands of novel or uncharacterized events (novel splice junctions, sequence variants, chimeric transcripts, noncoding RNAs, gene fusions, RNA editing events, etc.). Using genomics and transcriptomics data, one can then build a custom protein sequence database containing, in addition to known sequences taken from a reference sequence database, many predicted, novel sequences. By matching mass spectrometry data against this custom database one can obtain sample specific, protein-level evidence of expression for a subset of those events, likely bringing the biological or clinical significance of those events to a higher level. Thomas Kislinger: In combination, the coding proteome and noncoding transcriptome represent the end products of the sequence-to-phenotype continuum (DNA to RNA to Protein). The emerging view is that proteomic and transcriptomic approaches provide complementary readouts of the cellular state with neither holding a monopoly over defining the molecular phenotype. Of course, proteomics comes with the technical caveat that current technologies are not sensitive enough to identify every expressed protein sequence (at least the problem is more pronounced than in transcriptomics). Therefore, by combining genomic, transcriptomic and proteomic technologies, in a proteogenomic workflow, these technologies can inform each other. Classically, proteogenomics provides definitive protein-centric proof for the expression of a given DNA or RNA sequence (and perhaps more importantly, the mutant variants of these sequences). This peptide centric evidence can help refine current gene models and improve current reference protein sequence databases. As proteogenomics evolves beyond simply validating genomic predictions about the proteome, we will surely discover that the parameters perturbed to generate these predictions in the first-place can be reoptimized using evidence based on proteomic detection. Jacob Jaffe: Proteogenomics is both a complement to genomics and new paradigm for interpretation and visualization of proteomics data. In an increasingly genomics-centered world, proteomics (as a field) does itself a service by putting its data onto a scaffold that genomics folks can easily understand. Meanwhile, the underlying proteomic data can reveal things about biology that are inaccessible to genomics. How powerful is it to see that a phosphorylation site is recurrently mutated in certain cancers? That's the power of proteogenomics. Andrei Drabovich: Proteogenomics facilitates confirmation and correction of existing genes or even identification of potentially new genes. Current standard proteomic platforms rely on the reference genomic sequences and thus miss polymorphisms, mutated proteins, and rare protein variants. With multiple mechanisms leading to such rare variants, I would highlight peptides expressed by pseudogenes and noncoding RNAs as the most exciting area in proteogenomics, with the potential to identify some rare and even novel biological mechanisms. Proteogenomic data may also offer more reliable prioritization of cancer driver genes compared to genomic platforms, as was recently demonstrated for colon cancer. What type of variant peptides can be detected by proteogenomics (which are currently missed by classical proteomics)? Alexey Nesvizhskii: There is very long list of novel peptides that can potentially be identified using proteogenomics. These include novel splice junctions, peptides containing single amino acids variants, and peptides corresponding to alternative start sites. Other rare events include RNA editing events and gene fusions. In principle, one can detect peptides mapping to intergenic regions suggesting novel open reading frames, or to regions currently annotated as pseudogenes and noncoding RNA. Another category of novel peptides is peptides mapping to known protein-coding regions, including to their untranslated regions (UTRs), but in an alternative frame [e.g., peptides derived from upstream open reading frames (uORFs)]. Many recent studies reported identifications of all sorts of novel peptides mentioned above. Unfortunately, most of those studies did not apply the level of stringency in filtering their data that is required for detection of low-likelihood events (in most cases, the same filtering criteria were applied to the detection of known and as well as novel peptides). Thus, in my view, many of the previous claims of identification of rare events in published proteogenomics studies, including recent high-profile Nature studies describing the first draft of the human proteome, need to be critically evaluated. Thomas Kislinger: In theory any peptide sequence that is not present in a reference protein sequence database predicted from genomic/transcriptomic data and expressed abundantly enough to be detected by a modern shotgun proteomics strategy could be detected in a proteogenomics approach. This would include peptides with single-nucleotide variants (SNVs) and single-nucleotide polymorphisms (SNPs), insertions and deletions both in and out of frame, peptides that arise from aberrant splicing, and peptides that result from novel gene fusions. In addition, peptides that arise from translation of lncRNAs (long noncoding RNAs) or from intra- and intergenic regions of the genome could be identified by a proteogenomics approach. An additional caveat, aside from peptide concentration, might also be that some peptides are simply not amendable to mass spectrometric detection (i.e., biophysical properties). Jacob Jaffe: In the future it should be the norm that every sample analyzed by proteomics is in reference to the genome (or genomes) of the biological sample being interrogated. It's not a question of what is missed by “classical” proteomics; it's just that we're doing a bad job in proteomics by performing our analyses in reference to an average predicted proteome that is not really suitable for most samples. Andrei Drabovich: Such variant peptides include SNVs and missense mutations, fusion genes, truncated proteins, splicing isoforms, and peptides produced through translation of pseudogenes, UTRs, intergenic regions, or noncoding RNAs. Which biological or clinical unmet needs can be addressed by proteogenomic technologies? Alexey Nesvizhskii: Proteogenomics analysis can be useful as part of any study where more complete characterization of the genomic and proteomic diversity is desired. It is well established that that joint analysis of protein and mRNA data can provide biological insights not apparent from the analysis of each data type alone. Proteogenomics adds more depth to such integrative analyses by allowing detection and quantification of novel peptides that are expressed in a particular sample but would be missed when using a standard reference protein sequence database. From a proteomics perspective, for any biological or clinical question that can be addressed using proteomics and where relevant genomic data can be obtained (e.g., generated as part of that study or obtained from public sources), proteogenomics can provide an additional dimension. From the genomics/transcriptomics perspective, proteomics data can be extremely useful as a “proteomic filter,” suggesting which of the genomics-based findings are more likely to be functionally significant because they propagate to the protein level. Thomas Kislinger: Proteogenomics will continue to have an impact on genome annotation by providing direct evidence of what genes are translated to a protein From a biology of view, the integration of genomic, and proteomic data will in our of information from gene to Of course, such data will to a better biological or the identification of better clinical biomarkers is currently still the field is still in its It will on what of additional information are by using a proteogenomics example, it to that the of a mutant peptide (and thereby its compared to its could impact or to an unmet need is the characterization of peptides for important cancer genes that are suitable for for by multiple spectrometry and the of mutated to peptides in cancer can improve Jacob Jaffe: Proteogenomics really sets the for integrative biological a long has really just been a but by the genomics and the proteomics we can integrative analyses in Andrei Drabovich: biological unmet proteogenomic technologies may discover rare translational events in the such as expression products of pseudogenes and noncoding protein and reveal impact of missense mutations, thus providing a for clinical needs to be addressed by proteogenomics would include development of approaches using genomic and proteomic databases. This will facilitate more accurate of cancer and may to more What is the potential impact of in cancer Alexey Nesvizhskii: Proteogenomics is very relevant to cancer published and e.g., the by The genomics and transcriptomics technologies to generate profiles using cancer tissues as well as using generated as part of these studies contain many sequence variants and novel transcripts that are potentially important for the biological mechanisms of cancer or can be used as biomarkers for clinical An number of cancer studies include proteomic by the efforts of the that proteomic profiling of samples. it should be that the available proteomics and genomics data sets that are suitable for proteogenomics analysis are coming from Thomas Kislinger: the impact of on cancer is currently The and impact of in cancer will on many additional peptide sequences (and what can be identified by custom proteogenomics databases. what additional biological/clinical information (or can be obtained through the identification of such peptide The most is to the identification (and of an peptide can as a more sensitive can provide additional information for the of better or can be used to novel biological example, most genomic are likely to be mutations, could in reducing the by on that are and at the protein level. Jacob Jaffe: I will allow for between cellular and other and underlying processes. Andrei Drabovich: may complement genomics for of cancer discover molecular events upstream and of known cancer and the role of gene mutations, thus providing a for I would also that some clinical based on the proteomic analyses of mutated peptides might their in of rare and complement the next generation DNA of have a single missense in the gene where or These can be by the proteomic analysis of corresponding peptides is a proteomic could mutated at the depth of of one cancer detected in the presence of one in the the deep next generation DNA sequencing currently provides the depth of proteomic analysis of mutated peptides may facilitate the of some based on could of mutated peptides and thus diagnostic at diagnostic What are the technical in current proteogenomic Alexey Nesvizhskii: Proteogenomics high-quality proteomics data, generated in parallel with transcriptomic data. This is not and the number of deep proteomic data sets suitable for proteogenomics studies is still The is peptide identification is a and of proteogenomics. Many published including high-profile in Nature and other did not apply false discovery rate methods suitable for proteogenomics. In the stringency of filtering of novel peptides should be higher than that for commonly peptides. I have recently a of data analysis for proteogenomics studies, and we are to in this Thomas Kislinger: it is certain that omics technologies have not yet their and can still be I that the of proteogenomics are currently at the level of data This includes the development of appropriate statistical analysis to to the identification of proteogenomics peptides. In addition, one could even that we even what is the most appropriate proteogenomics example, is an approach combining NGS and shotgun proteomics on the same samples the approach as a what type of sequencing, exome-Seq, or a customized database using available variant (or peptide sequences Of In could be the of my knowledge this has not been to and each strategy has its and to be addressed by proteogenomics is with to data like genomics, proteomics now has the to detect in combination, have the potential to identify the This potential with to the of raw data that will have to be by the Jacob Jaffe: need to it routine to a customized proteomics database from any type of genomics/transcriptomics data. need to be to of in the of these proteogenomic The between gene and Andrei Drabovich: of mass needs to be by an additional to to enable the quantitative analysis of proteins the of of of protein in and biological In addition, software for proteogenomic analysis and better statistical to false-positive and false-negative of peptide matching will be In what the field of proteogenomics to in the Alexey Nesvizhskii: Proteogenomics has been for but its impact has been in part to depth and protein sequence coverage in data produced by the previous of proteomics However, has and now allows deep proteome profiling, in some the depth of RNA-Seq data. one can start using proteogenomics, in a more than to for evidence of protein-level expression of transcripts currently annotated as noncoding RNAs or pseudogenes, to search for and The number of transcripts in any RNA-Seq data for any sample is the question on is which of these are Which have any There is a of interest in noncoding RNAs and with about their protein coding is a of as Proteogenomics can provide valuable information As ribosomal profiling technology the combination of RNA-Seq, ribosomal profiling, and proteomics will provide a of data for all sorts of proteogenomics Thomas Kislinger: the have been addressed and peptide are including the of the most appropriate proteogenomics one is that most cancer proteomics will apply some type of proteogenomic approach. The next would be to detection of aberrant peptides and to the of such proteins in the of cancer biology or example, does the expression of a mutated protein its protein or In the of one could proteomics (i.e., to peptides. additional could be to include top-down with its to specific in a proteogenomics this might not be technically as of this could provide the of the cancer proteome is a of the upstream cancer Jacob Jaffe: at a proteogenomics in a genome quantitative information and will be will be as more and more data Andrei Drabovich: will be integration of genomics, transcriptomics, and proteomics by not but also quantitative Such integration will be by protein and There is also a hope that proteogenomics will better of very exciting to the molecular of cancer could be for some rare example, neither gene were identified for type for which cancer driver may at the level of proteome or I also that proteomic will start using the genomic data generated by the large genomic such as the and example, the contains comprehensive genomic data, such as the whole exome sequencing, profiling of and DNA and mRNA and for more than these data will be translated into proteomic databases and of proteogenomics. UniProt knowledge base next generation sequencing tandem mass spectrometry whole genome sequencing untranslated upstream open reading frames single-nucleotide variant single-nucleotide polymorphisms multiple spectrometry The false discovery rate
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,107 | 0,173 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,003 |
| Bibliométrie | 0,004 | 0,004 |
| Études des sciences et des technologies | 0,003 | 0,017 |
| Communication savante | 0,008 | 0,017 |
| Science ouverte | 0,007 | 0,008 |
| Intégrité de la recherche | 0,005 | 0,020 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,005 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».