Whole Transcriptome Profiling: An RNA‐Seq Primer and Implications for Pharmacogenomics Research
Bibliographic record
Abstract
Pharmacogenomics has revealed compelling genetic signals associated with variability in drug response. Gene expression studies represent an additional approach to identify candidate genes accounting for drug response variability. This review focuses on insights that might be gained through analysis of the transcriptome to reveal the influence of gene expression on variable drug response. We provide a basic overview of RNA-Sequencing (RNA-Seq) and its applications, and outline advances in pharmacogenomics achievable with RNA-Seq data. Every human cell in the body arises from the same set of genetic information, yet only a fraction of genes is expressed in any given cell at any given time.1 This carefully controlled pattern of gene expression differentiates liver cells from muscle cells, for instance, and healthy from diseased status. Therefore, enhanced understanding of gene expression patterns can lead to molecular pathways that underlie disease susceptibility or drug response. The complete transcriptome consists of protein coding and long and short noncoding RNAs. We will focus here on protein coding RNAs (mRNAs). The expression level of RNAs represents the most immediate phenotype that can be associated with cellular conditions (such as drug exposure or disease state), and regulatory variants in the gene locus itself (cis-acting) or in trans-acting regulatory factors. Sequence variation in regulatory regions that govern gene expression is a main mediator of overall phenotypic diversity.2, 3 On the other hand, genetic variants in the transcribed region of a gene can influence multiple RNA functions, such as splicing, turnover, and translation.4 Therefore, RNA levels reflect the combined influence of genetic factors, cellular conditions, and environmental factors. We propose that regulatory variants are key factors and frequently represent causal mutations in disease genetics5 and pharmacogenomics.6 High-throughput DNA sequencing tools have provided a new, comprehensive method for both mapping and quantifying transcriptomes.7 RNA-Seq has emerged as an innovative method for both mapping and quantifying transcriptome signatures associated with diseases and traits.8-10 When compared with other transcriptomic techniques, such as microarrays, RNA-Seq has the ability to quantify expression levels of all RNAs at a given gene locus, including RNA isoforms generated through alternative transcription and translation start sites, 3’UTR poly-adenylation sites, splicing, RNA editing, and more. As a result, RNA-Seq characterizes the complete transcriptome and facilitates discovery of differentially expressed genes and RNA isoforms that are not otherwise accessible. The power of sequencing RNA vs. using oligonucleotides to assess gene expression with microarrays lies in the fact that both transcript discovery and quantification can be incorporated in one high-throughput sequencing assay with RNA-Seq. Thus RNA-Seq enables dynamic assessment of mechanisms associated with disease and drug response to bridge the gap between genomics and phenotype,11, 12 providing a powerful tool germane to precision medicine. Until recently, microarrays have served as the most cost-effective, reliable, and rapid technology for high-throughput profiling of gene expression. However, microarrays require a priori knowledge of sequences to be investigated, limiting discovery of de novo splicing isoforms or novel exons, transcripts, and genes (Table 1).7 In addition, hybridization-based methods used in microarrays can also limit the dynamic range of gene expression quantification (Table 1), casting doubt on measurements of transcripts with high or very low abundance.13 With widespread adoption of Next Generation Sequencing (NGS) platforms, RNA-Seq, a methodology for RNA profiling, using millions of short reads (sequence strings), enables the investigation of all the RNA in a sample, theoretically.14 In practice, the input population of RNA, either total RNA or fractioned (such as poly(A) selected, capturing most mRNAs and many noncoding RNAs), is converted to a library of fragmented cDNA.14 Then, each fragment receives adaptors attached to one or both ends.14 These fragments are amplified and sequenced in a high-throughput manner, generating millions of short reads14 (Figure 1). Current RNA-Seq methods target RNAs with at least 200 base pairs, whereas short noncoding RNAs, including microRNAs, require separate isolation and protocols.15, 16 Depending on the sequencing platform (Illumina, Roche 454, Solid, Ion Torrent), read lengths typically range between 30–500 base pairs.17 Sequence length is an important criterion since longer reads improve mappability for identification of transcript and transcript isoforms.18 Another important factor is the library size or read depth, which is the number of sequence reads for a given sample. The deeper the sequencing level, the more sensitive and precise transcript identification and quantitation will be.18 While some studies advocate that read counts as low 5 million reads can accurately quantify moderate to highly expressed genes,19 the ENCODE best practices protocol recommends library sizes with more than 25 million reads for a typical RNA-Seq protocol for investigating mRNA expression using poly-A selected RNAs.20 Once high-quality reads are obtained, RNA-Seq reads are computationally mapped to the human reference genome, revealing a transcriptional map.7, 21 Owing to extensive alternative splicing that occurs in the human transcriptome, the alignment process is more challenging to map reads that span splice junctions.17 Also, RNA-Seq read alignment is complicated by the fact that short reads may be assigned to multiple regions of the human genome.17 The most widely used RNA-Seq alignment software programs use gene annotation to achieve better placement of spliced reads and correctly handle multiple short read assignment in the vast majority of occurrences.22 Next, overlapping reads that were mapped to a particular exon are clustered into gene or isoform levels for quantification.18 Raw read counts per gene locus alone are insufficient to compare expression levels among samples.18 The most frequently reported measure of gene expression from RNA-Seq analysis is R/FPKM (reads or fragments per kilobase of exon model per million), a within-sample normalization method that considers transcript length and total number of mapped reads.18 The data analysis then allows the characterization of gene expression levels that can be applied to investigate distinct features of the transcriptome diversity. As with all large-scale analyses, the resulting RNA levels are subject to error, so that important findings need to be replicated with alternative methods, such as quantitative real time-polymerase chain reaction (qRT-PCR). The beauty of the RNA-Seq tool lies in the fact that previously distinct core activities of discovery and transcript quantification now can be combined in a single high-throughput assay. This approach provides a significant qualitative and quantitative improvement to study the transcriptome, enabling detection of genes with low expression (given enough read counts per sample), sense and antisense transcripts, RNA edits, and novel isoforms, all at base pair resolution.7 One of the most biologically relevant applications of RNA-Seq is the comparison of mRNA transcriptomes across distinct developmental stages, across samples from diseased vs. normal individuals, or other specific experimental conditions.23 For this type of analysis, it is crucial to accurately construct the isoform structure to assess transcript abundances when comparing multiple samples (Figure 2).17 This powerful approach is essential for the interpretation of functional genomic elements and discovery of transcripts key to molecular mechanisms underlying disease susceptibility or drug response. Alternative splicing events play a key role in shaping biological complexity and genomic diversity.24 The term alternative splicing refers to distinct inclusion/exclusion of exons in the processed RNA product when compared with constitutive splicing events.25 Multiple proteins regulate this RNA processing step, called splicing factors, aggregated into tissue-specific spliceosomes.25 Given the complexity of this regulatory activity, it is not surprising that RNA splicing is exceptionally susceptible to hereditary and somatic mutations associated with a broad range of diseases.24, 26, 27 The RNA-Seq technology enables the exploration of transcriptome structure, investigating different patterns of splice junctions with more accuracy than microarrays.28 Once sufficient RNA-Seq reads (tens to hundreds of millions) are mapped to the genome, exons, and exon–exon junctions, RNA-Seq assays allow annotation of new exon–intron structures and detection of the relative isoform abundance of individual alternative splicing events.28, 29 Unlike microarrays, RNA-Seq does not rely on prior knowledge of transcriptome structure and splicing events, and has nucleotide resolution level.30 Deep surveying of alternative splicing with RNA-Seq data has revealed unprecedented diversity of splice junctions, tissue-specific RNA-binding motifs, and splicing regulatory elements.31 The relevance of alternative splicing is further highlighted by distinct functions conveyed by splice variants that can contribute to tissue-specific pathology.5 Most of the single-nucleotide polymorphisms (SNPs) identified through genome-wide association studies (GWAS) reside in noncoding or intergenic regions of the genome,32 suggesting that many causal variants influence traits/phenotypes by impacting gene expression.33-35 Genetic polymorphisms associated with variation in gene expression levels, termed expression quantitative trait loci (eQTLs), have been extensively studied over the years and are known to be widespread over human populations.35, 36 These regulatory variants contribute to phenotype diversity by interfering with the steps across the flow of genetic information in a cell, from DNA to protein, and are cataloged now on GTEx37 for up to 60 different tissues in up to 1,000 subjects. RNA-Seq enables further investigation of the regulatory role of specific sequences to gene expression by taking advantage of the single-nucleotide level resolution. Heterozygous individuals for a particular genome locus present two allelic forms, which allows one to investigate whether one of the alleles has greater expression than the other. This event is called allele-specific expression (ASE), and detection of ASE imbalance with RNA-Seq (unequal expression of one allele over the other) signals the presence of genetic and/or epigenetic determinants that govern allelic transcriptional activity (Figure 3).38, 39 Often, ASE is evidence of a disruption of a highly regulated process leading to disease susceptibility38, 39 or potential variability in drug response.40 Predominantly, the largest effect sizes or the strongest genetic effects in the expression of individual genes are observed locally within the respective target gene locus.36, 41 These are called cis-regulatory regions, composed of cis-regulatory elements, with target sites for transcription factors and other regulatory proteins, acting as promoters and enhancers, or as repressors defining transcriptionally inactive regions.36 Transcription factor binding sites are the central elements of cis-regulatory regions, which in the presence of transcription factors and epigenetic modifications can determine whether transcription is turned on/off, and the rate of the transcription process.36, 42 Enhancer regions can reside at large distances up- or downstream of the gene locus per se, often confounding the assignment of GWAS hits to a candidate gene.43 Trans-acting variants, polymorphic variants that regulate gene expression via an intermediate factor, can be anywhere in the human genome, and typically convey a smaller-effect size than cis-acting variants.33, 36, 42 One of the reasons may be that expression levels of a particular gene are usually under the effect of multiple trans-acting regulators, such as different transcription factors, coactivator proteins, proteins that help stabilize transcription factors, etc. Consequently, the effect size of each one of these trans-acting regulators is diminished.42, 44 To date, several trans-acting regulatory regions have been identified as “hot spots” but only a few of these regions have been determined to account for the underlying regulatory mechanism.44-50 RNA-Seq data can be further explored to infer gene function, gene–disease or gene–drug exposure associations and gene–gene interaction with coexpression network analysis, an approach that constructs networks of coregulated genes.51 Going beyond the identification of singular genes or regulatory variants associated with disease or drug exposure, building coexpression networks can be used for candidate gene prioritization as a function of their position in network hubs, and functional gene annotation.52 Guidance and further details about the various methods developed for this approach are available in recent reviews.53, 52 Because RNA-Seq also quantifies the expression of up to 70,000 noncoding RNAs,54 not usually measured with microarrays, it permits a better understanding of regulatory networks driving biological processes including noncoding RNAs. Numerous noncoding RNAs are thought to have regulatory roles55 and to play a role in disease processes.56, 57 With sufficient read depth, RNA-Seq also increases accuracy for low abundance transcripts18 and has the requisite resolution that allows to distinguish between the expression of different splice variants.58 Thus, coexpression analysis on RNA-Seq data can detect previously hidden networks and thereby assign putative functions to noncoding RNAs and splice variants. In the past decade, GWAS have been the most widely employed tool to investigate the link between genetic polymorphisms and common diseases, due to the application of agnostic approaches in which genetic variation across the human genome is tested, allowing discovery of novel genes and pathways. Although this approach has successfully revealed a multitude of genetics signals associated with complex diseases and phenotypes, revealing new biological insights, often gene expression studies applying RNA-Seq identify signature genes that explain a greater fraction of interindividual variability than large GWAS. Recently, a sequential series of large-scale GWAS in hypertension (HTN) has been published. Data from two meta-analysis studies, each with greater than 30,000 individuals, identified three replicated loci in association with HTN.59, 60 A few years later, a study with an even larger sample size identified 29 loci with significant associations with loci been previously associated with These were with in the genetics new insights into the of However, each only effect about per allele for and per allele for which in for than of interindividual variability in Therefore, these studies have not provided signals for defining of and further exploration of the mechanisms highlighted by these genetics findings is RNA-Seq approaches have also been used to understanding of A investigation of gene expression signature using RNA revealed genes that in explain up to of interindividual variability in These on exploration of expression in to of variability in by the GWAS findings of the signature and GWAS revealed that associated with 5 in the are also regulators of several signature Therefore, this study provides important for investigation on the of these transcriptomic to drug and as an of the insights to be gained from RNA that are and amplified when with GWAS data. the application of RNA-Seq in for transcriptome profiling revealed novel potential mechanisms in the of and its identified genes and biological pathways associated with a effect on identified genes of for through transcriptome characterization of the of under conditions and expression and analysis revealed genes in as a potential in Multiple recent studies have also the gap between human regulatory variants, gene and One is the gained on the and variants associated with This region to as an with the gene which is more than from the variants, its gene expression in both and human a causal link between and a large-scale study with RNA-Seq data from the a genome-wide for on the of gene expression in multiple tissues and cell This study identified 16 cis-acting regulatory variants and one trans-acting the expression of genes in the of investigating the role of in downstream the RNA-Seq and effects that on and identified of gene expression variants that to the potential for new and provide a better assessment of individuals at for Another study of gene expression with RNA-Seq that transcriptome profiling when compared with GWAS data. With a that allows data from multiple and different effects between these a recent study data from the for of of on gene expression profiling with on used for the assessment of the that gene expression data provide more power than any assessment in the and the of gene expression and provided a significant in These the power of gene expression studies to or stages, and information can be with RNA-Seq data. In recent multiple studies have the transcriptome signature of expression analysis comparing transcriptome between human and human and were identified as potential in human The same also identified long noncoding RNA differentially expressed between normal vs. Another study used generated by RNA-Seq and microarrays, to identify novel gene expression signatures of Although these findings are not yet for provide a comprehensive characterization of the transcriptome in human and represent an of key in for further investigation with transcriptome can to for key genes in for RNA-Seq data for of in the and These were used to the whether dynamic between genes at the gene or transcript levels can to identify key factors in disease This approach to the discovery of several networks association with suggesting that single genes and transcripts alone are insufficient to account for disease but dynamic need to be via analysis of RNA-Seq data. The provided here from the in disease genomics are not an extensive of each study available with RNA-Seq data in the but the provided by these studies set the for studies in pharmacogenomics and some of the potential for discovery by RNA-Seq data. have that gene expression studies, with RNA-Seq, account for greater variability in disease and that including gene expression data have the accuracy of disease Pharmacogenomics the to provide taking into the genetic Gene expression variation and the diversity of splicing events in drug and drug have been associated with in drug response and drug Therefore, a comprehensive study of the potential variability on transcriptome profiling associated with pharmacogenomics can provide relevant insights into the molecular of in drug response. In this provide some of pharmacogenomics studies that have applied RNA-Seq technology for identification of have transcriptome sequencing for the investigation of disease genetics the However, the use of technology for pharmacogenomics has been Given the potential of a study of the transcriptome to drug response the of Pharmacogenomics the for a comprehensive transcriptome sequencing that has cataloged variation in gene expression and splicing events of in drug across and cell Gene expression and splicing data are available for data from cell enables pharmacogenomics to and insights, and to identify of drug The RNA-Seq and genetic number and data from 1,000 cell with for to data and revealed and of drug highlighted a few cell for expression for and expression for a to the gene expression analysis of the candidate genes associated with drug response identified and other or cell genes as key functional for the association with drug response to the The study also revealed distinct coexpression patterns of drug response between and These provide new molecular and networks to RNA-Seq data are available for profiling in tissues has also in the identification of drug response in of genes associated with molecular in The presence of this gene expression signature potential with a The same other molecular signatures to on an extensive functional and genomics This the for that can from gene expression investigation with RNA-Seq data. for with is an of a number of that have only for a of with specific molecular RNA-Seq for an transcriptomic approach from and a set of coding and noncoding associated with with cell functional investigation of candidate that of in an gene expression investigation with RNA-Seq, this study revealed relevant of and potential for novel analysis in cell from of the and to identify genes that may have a role in the genes with gene the most relevant biological with extensive that this gene to the of The also a of with response and that expression levels and splicing for about of the variation on response in applied RNA-Seq to investigate a effects of for the of muscle cells were with for by mRNA and high-throughput This approach identified genes differentially expressed relative to and highlighted as an candidate gene that effects of of transcriptome and drug response are under in from the RNA-Seq the study of gene expression impacting response. In to identify novel molecular of response to from the Pharmacogenomics of and studies with of response and to from and from were selected for this RNA-Seq data set to assess the gene expression levels of genes previously associated with expression relative to revealed that and were differentially expressed in all three These findings that genes identified through transcriptome profiling are also relevant determinants of response to In addition, the RNA-Seq data provided further biological insights when with the GWAS The associated with better response to and and with expression of the RNA-Seq data analysis revealed that expression of in as to when compared with and The allele-specific expression analysis also revealed a expression imbalance at which the observed genetic effects through expression and/or studies are to identify the potential of RNA-Seq data in understanding variable to We can to the in the term of data from transcriptome analyses, understanding of mechanisms and of interindividual in drug response. The application of RNA-Seq may lead not only to the discovery of signature genes of response to but it may also the characterization of isoform regulatory variants, and gene expression networks impacting in drug response. This powerful tool represents an alternative approach for the identification of target allowing a of RNA transcripts or transcript in the mechanisms underlying drug response. Sequencing have in the past allowing the application of RNA-Seq to investigate the transcriptome with accuracy and high data resolution to the knowledge on the influence of regulatory mechanisms on gene expression variability in drug response. In this review this new technology and multiple applications for RNA-Seq, with of in disease genomics and studies are to the pharmacogenomics to the level of knowledge enabling the findings from novel genetic variants associated with variability in drug studies have potential to While the use of in pharmacogenomics is recent advances in allow transcript quantitation for expression between biological conditions, identification of splicing events, and the assessment of regulatory mechanisms of gene expression. These are processes generating diversity in function with in drug of and has into the of challenging to the methodology for With the of new powerful methods, the study of the transcriptome, over the is to of disease susceptibility and drug This in by from the of Pharmacogenomics and and and the The of
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".