Perspective: How to Make Microarray, Serial Analysis of Gene Expression, and Proteomic Relevant to Day-to-Day Endocrine Problems and Physiological Systems
Notice bibliographique
Résumé
The recent mapping of the human genome has opened the way to novel technologies for identifying new genes or clustering regulated genes and proteins in a tissue- and cell-specific manner. Microarray and serial analysis of gene expression (SAGE) are quite powerful in this regard, and they generate on a day-to-day basis data banks of gene profile from cells or tissues of animals challenged with different compounds including hormones. Here we review these techniques with their strengths and limits in physiological systems with a particular emphasis on the endocrine system. We also present new developing techniques in proteomic, especially for the analysis of functional proteins and protein-protein interaction. Although we are still at the embryonic stage of the proteomic area, there is no doubt that the current ongoing work will have a great impact in endocrinology. There are also serious drawbacks that have to be taken into consideration to prevent generating data that may either be nonphysiological or difficult to interpret, the most critical one being the experimental design. Since the crystallization of the first hormone—adrenalin—by Takamine and Aldrich at the beginning of the twentieth century, modern endocrinology kept growing to comprise various fields of research, namely cytology, cellular biology, and more recently molecular biology and genetic. During the past two decades, large-scale sequencing efforts including the Human Genome Project (1) generated initial large-scale databases, and numerous new genes were discovered, some of them encoding proteins with a function still remaining to be unraveled. Recent progress in biotechnology, more particularly in gene expression microarray and SAGE technologies gave us new tools for identifying gene functions and much more (see Table 1). Actually, microarray and SAGE experiments allow us to test the expression of thousands of genes simultaneously and to identify automatically the genes of interest. In the same way, based on the last developments in the technologies of protein separation, quantification, and identification, protein expression profiles are now available with proteomics. Because proteins are final posttranslational products from mRNA, proteomics will give us access to a new database with particular biological significance. Today, in the scientific literature, the number and diversity of data generated from microarray experiments are impressive and there are already numerous reports covering the whole biomedical community. In the field of endocrinology, analysis of both gene and protein expression will be quite powerful tools to study the regulation of physiological mechanisms triggered or inhibited by hormones. Here we review how to make DNA microarrays, SAGE, and proteomics relevant in modern endocrinology. List of various studies in the field of endocrinology that exploited DNA microarray or SAGE technologies List of various studies in the field of endocrinology that exploited DNA microarray or SAGE technologies DNA microarrays have been developed to exploit the huge amount of sequence data generated by large-scale sequencing programs. Briefly, fluorescent probes prepared from the mRNAs of the samples are hybridized onto high-density matrix of thousands or tens of thousands of ordered DNA known sequences representing specific genes. Hybridization intensity is determined for each represented gene on the matrix, allowing the quantitative comparison of the expression levels of almost all transcripted genes in two or more RNA samples. Two different microarray technologies are available; the oligo microarrays (e.g. from Affymetrix, Inc., Santa Clara, CA) and the cDNA microarrays that differ with the length of DNA sequences (from 25 oligomers to several hundred oligomers, respectively) synthesized or grafted on the matrix, the type of the matrix (glass, nylon, membranes, and other formats) and, finally, the data processing. Arrays are customizable in DNA species and in number of genes represented. When using two different samples (treated and control), we can compare the gene expression profiles between them and then determinate how the cell or tissue regulates its genes in a specific environment. DNA microarrays are like powerful automatic RNA differential display experiments, without the need to both sequence and quantify the bands of interest. Moreover, cDNA microarray sensitivity allows working with as few as 10 μg RNA, for instance about 100,000 cells (2), which is compliant with the small quantities of clinical samples needed in endocrinology. Thus, DNA microarrays are suitable tools for endocrinology studies, such as the analysis of the cellular response to a specific stimulus. For example, Feng et al. (3) identified from mouse livers 45 genes not previously identified as thyroid hormone-responsive genes. In another example, Dupont et al. (4) have used cDNA microarray technology to define the specificity of insulin vs. IGF-1 signaling. Of the 2221 genes tested on cDNA microarrays, 30 genes significantly increased in presence of IGF-1 but not by insulin, and 27 of them were not previously reported as being IGF-1-responsive genes. Work done in other fields can also be quite powerful to unravel genes associated with the endocrine system not necessarily expected to be regulated or even expressed in a particular group of cells. In exploring how dendritic cells modulate the immune system in response to different pathogens, Huang et al. (5) found that activinβa is one of the highly up-regulated genes when antigen-presenting cells are exposed to Escherichia coli. It will not be long before we clarify the physiological relevance of such regulation and the presence of this hormone in the innate immune system. One may expect unraveling very fine and unexpected mechanisms modulated by activin in the innate immunity, which is the case for SMAD3 that is one of the intracellular mediators of the activin receptors (6). As a major application of DNA microarray, expression arrays can be used to understand multigenic diseases such as many cancers (7, 8). “Fold difference,” e.g. the ratio of gene expression in a treated sample over the control sample is used as quantitative measurements of the differential expression, to generate a clustering of genes. These clusters can be arranged hierarchically or spatially to form self-organized maps (9, 10). Expression cluster can be used to search common motifs of genetic regulation to find new regulatory mechanisms. Another application of DNA microarrays is the finding of new functions of genes by association of gene expression. Bioinformatics, with the use of algorithms, provided tools to trace out metabolic pathways, cellular interactions and to discern genetic networks (11–13). However, before using these predicting algorithms, it is imperative to distinguish between significant fold difference values and false-positive results to avoid reporting data that may actually not be physiologically relevant (14). Replicates are also required to lower the experimental noise and to display low level of differential expression significantly. For these reasons, it is imperative to confirm the presence of newly discovered expressed genes in a specific tissue by Northern blot, real time RT-PCR or in situ hybridization. The latter being the best approach, because it permits not only to validate, but also visualize the expression pattern and even the type(s) of cells expressing the transcript during a specific time, treatment, or changes in plasma hormone levels. As described above, DNA microarrays are useful to identify genes that are markers of multigenic diseases. When these markers will be well defined, the design of DNA microarrays could then be customized to test simultaneously all these markers for a diagnostic use. Furthermore, DNA microarrays will be very efficient tools to detect the response to therapy, such as the prostate tumor response to androgen withdrawal, and to plan more appropriated medical treatments. In this regard, Bubendorf et al. (15, 16) have described another use of microarrays, not as DNA microarrays, but as tissue microarrays. Concisely, hundreds or thousands of 0.6-mm diameter tissue cylinders are arrayed on a glass slide allowing the instantaneous analysis of every sample with either immunohistochemistry, fluorescence in situ hybridization or RNA in situ hybridization. Therefore, these tissue microarrays could be quite useful for clinical studies, such as paired analysis of prostate cancer biopsies. When DNA microarrays are used in a whole-genome expression analysis, large volumes of data are generated raising computational requirements (17). Many microarray data are now available on public on-line databases (e.g. Stanford Microarray Database; http://genome-www5.Stanford.EDU/MicroArray/SMD/). However, analysis of DNA microarray data are limited to relative comparison between samples. Furthermore, oligo microarray raw original data have to be processed for bias corrections like multiplicative effects (e.g. difference in the total mRNA concentration of samples), additive effects (e.g. background), position effects on the microarray and nonlinear effects (saturation of detectors of the hybridization intensity). All these biases emphasize the need to have access to complete raw data sets to provide a significant comparison of array results when processing data with normalization curves. Moreover, the identification number of each gene on the array is usually different from one microarray brand to another one, requiring the usage of Unigene nomenclature to share and compare different platform microarray data (18, 19). Expression profiling of large amount of clinical samples is very efficient with microarrays, although SAGE is more suitable than microarrays for identifying new genes or RNA that are alternatively spliced, because microarrays allow only to test known genes on the chip. SAGE is a method based on the isolation of unique short sequence tags from individual polyadenylated RNAs and on concatenation of these tags serially to facilitate their sequencing, and therefore examine gene expression profiling (20). Polyadenylated RNAs are captured from cell lysates with oligo-deoxythymidine-coated beads and are reversed transcripted in cDNA. Isolation of tags from cDNA is performed with the formation of unique 5′ end position within the 3′ end part of each cDNA by cleavage with anchoring enzyme. Tags are released using different strategies (21) and are concatemerized into long DNA sequence. Finally, concatemer clones are sequenced and tag sequences are BLAST against GenBank to allocate a gene identity to each tag. Basic tag counting permits the determination of absolute tag abundance. There are several limitations to keep in mind when using SAGE. For example, some transcripts could lack an anchoring enzyme site and would not be tagged. There is also an inherent low sequencing error rate that alters the accuracy of the tag count and increase mistrust of the abundance of tags with low count. Another problem is the making of valid tag to gene assignments while the large majority of transcript source sequences available in GenBank are expressed sequenced tag sequences. These are usually only single-pass sequenced, making possible to contain sequence errors. Additionally, tags are very short sequences (usually 9–11 bp), and two genes can share the same tag. A further source of the problem is when making a tag-to-gene assignment for a tag without corresponding entries in databases. Because the sequence available in the 11-bp tag is extremely limited, the cloning of the full-length genes then becomes difficult. On the other hand, SAGE strengths are remarkable. First of all, SAGE data represent absolute RNA expression levels that are easily portable and directly comparable to existing SAGE database. Actually, more than three million transcript tags are already available on the Internet (http://bioinfo.amc.uva.nl/HTM-bin/index.cgi/; http://www.sagenet.org; http://www-dsv.cea.fr/thema/get/sade.html; http://www.ncbi.nlm.nih.gov/SAGE; http://www.urmc.rochester.edu/smd/crc/swindex.html; http://genome-www4.stanford.edu/cgi-bin/SGD/SAGE/querySAGE), and the number of libraries keeps growing. SAGE also allows the potential identification of new transcripts that are not already recorded in GenBank. SAGE is compliant to analyze the differential gene expression between diseased and normal tissues, and studies have been reported on diseases, such as arteriosclerosis (22) and human immunodeficiency virus infection (23). This technology has recently been used to identify the full set of genes expressed by mammalian rods that provided evidence that half of all cloned human retinal disease genes are selectively expressed in rod photoreceptors (24). SAGE has been widely used in the fields of immunology and neuroimmunology as well as oncology (25–28). For example, Polyak et al. (29) reviewed some applications of SAGE in cancer research and described more particularly the analysis of specific gene expression patterns in cancer cells and also the identification of regulatory targets of oncogenes and tumor suppressor genes. Some interesting applications of SAGE have also been reported in endocrinology, such as the changes in the transcriptome of kidney cortical collecting duct principal cell line induced by aldosterone and vasopressin (30). After sequencing approximately 170,000 transcript tags, roughly 15,000 tags were assigned to identified genes, whereas 3,642 tags failed to match with known mouse sequences. This work revealed 34 aldosterone-induced transcripts, 29 aldosterone-repressed transcripts, 48 vasopressin-induced transcripts, and 11 vasopressin-repressed transcripts, some of them having been validated by Northern blot hybridization or real-time RT-PCR (30). With a similar strategy, Datson et al. (31) reported the identification of over 200 putative corticosteroid-responsive genes in rat hippocampus that are regulated via mineralocorticosteroid and glucocorticosteroid receptors. These corticosteroid-responsive genes could provide new insights on the role of glucocorticoids in the brain and their potential involvement in the mechanisms leading to neuroprotection and/or neurodegeneration (for a review, see Ref. 32). Another example of SAGE application in endocrinology is the exposure of cancer prostate LNCaP cells to synthetic androgen, which resulted in 136 induced genes and 215 repressed genes when compared with untreated control cells (33). Most of these androgen-regulated genes were not previously described, underlying again the role of SAGE technology to discover new genes and their functions. Although a good correlation between transcript and protein expression levels is expected, mismatches can occur (34, 35), because posttranscriptional mechanisms control the turnover and the posttranslational modifications of proteins. Moreover, alternative splicing can generate multiple transcripts that enhance the diversity of protein functions. Thus, information about protein expression is both important and complementary to genomics, opening therefore the way to the proteomic area that is clearly under way at this time. Like genomics, proteomics take advantage of the later developments in high technology to allow, as initial goal, the mapping of the proteome of biological systems. One primary tool in proteomics is the protein separation by two-dimension gel electrophoresis (2DGE) (36) followed by immunoblotting or protein visualization with either a staining (Silver, Coomassie, or SYPRO Ruby) or other chemoluminescence or radiolabeling methods. 2DGE techniques provide the first protein fingerprints in a single picture proteins extracted from tissues or cells. Comparative picture analysis with computers then leads to the identification of differentially expressed proteins guiding their extraction from gel. One limitation of 2DGE fingerprints is the difficulty to compare and quantify low protein expression levels, but this problem could be bypassed using isotope-coded affinity tags (37). Until recently, proteins were mainly sequenced by Edman method, which was limited in sensitivity and restricted in N-terminal modifications. Actually, mass spectrometry (MS) seems the method of choice to characterize proteins (reviewed in Refs. 38 and 39). After digesting of extracted proteins with trypsin, peptide masses are commonly measured using either matrix-assisted of or and are compared with protein databases or databases to characterize the mass databases are already available via 2DGE associated with has been used in various studies to protein expression differential display but not in the field of endocrinology at time that this review was It not be long before reports using these powerful to the role of on protein profile in tissues and cells in In to these proteomic there are other tools that are under to the new of functional proteomics. Like cDNA microarrays, arrays based have been developed for of protein-protein interactions or Actually, there are already of for protein or peptide microarrays and arrays Moreover, have developed the which on proteins of the sample on a specific matrix or and to with a and the peptide masses are measured by These are useful to and clustering of although they not provide information on the changes that occur during protein limitations of protein arrays are both the of the grafted proteins and their in Furthermore, short peptide microarrays not take into the effects of the protein environment. As an alternative to protein microarrays, and have developed a microarray of cells expressing cells with the cDNA when they are on an ordered cDNA The microarray of the can be for some limitations at the level of protein these have been developed These and into two and can be to a mass via an can be used to protein as well as and proteins with a sensitivity and than separation by 2DGE and analysis can also be with the Briefly, this allows in the same time the and the of protein from the 2DGE to a that is using tools are clearly limited by the but they are and there is no doubt that we will to a proteomic in the few will from this that will finding new and small It will also be possible to provide functional in gene and protein expression data with other data such as automatic information extraction and DNA and protein sequence database. In modern endocrinology, DNA microarrays and SAGE can us to identify the hormone-responsive and proteomics will allow the identification of hormone-responsive (see 1). DNA microarray and SAGE as well as 2DGE are now tools in numerous progress will from computational more particularly from and also from the of all biological information databases genes, data In of this huge amount of information is a way to analyze an of both the transcriptome and how these can be for useful data from the gene to the physiological DNA microarrays, SAGE and proteomics in a of experimental tools and the The different levels of a biological system are represented in with the technology that can be experimental of the results and data All the are to provide information in an and of significant physiological DNA microarrays, SAGE and proteomics in a of experimental tools and the The different levels of a biological system are represented in with the technology that can be experimental of the results and data All the are to provide information in an and of significant physiological The and the large amount of data that are generated by these new technologies are the current for making them available on a day-to-day basis for Another that is when microarray, SAGE, and proteomic is the experimental design either in or in physiological is a data generated from tissues of animals treated with large of for example, than physiological of glucocorticoids that are during It is that the system will be to identify new genes or clusters of regulated genes and but such will be during normal endocrine changes an The may also be a potential especially in highly tissues and when take It is quite difficult to the group of regulated genes is by cells within the tissue or from immune cells. This is especially during of immune but also during normal most tissues are with and its In situ hybridization is an important for the data generated via either microarray or SAGE This has numerous including the pattern of expression and cellular source of the hybridized genes. identification is best in and the whole with Therefore, a lack of hybridization may not necessarily the microarray and SAGE because these regulated genes may be expressed by cells that are no present in tissues of One can this when gene analysis is performed on tissues, such as and a large cluster of regulated genes will most be of and not It is therefore quite important to take these into to avoid such mismatches and in the data The new technologies for gene and protein analysis will be quite for field of research, but one keep in mind that experimental design is the first and most important this is even the most and will be in the of data that will be On the other hand, small and experiments are to generate data of high physiological relevance that will be on a day-to-day a large-scale the need of numerous in almost of all fields of research, but will a role to make that all this is The of the of is an and a in gel matrix-assisted of mass serial analysis of gene expression.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,003 | 0,007 |
| Science ouverte | 0,003 | 0,001 |
| Intégrité de la recherche | 0,006 | 0,007 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,012 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».