MétaCan
Menu
Retour à la cohorte
Enregistrement W2124082515 · doi:10.1158/1055-9965.681.13.5

SNPs, Haplotypes, and Cancer: Applications in Molecular Epidemiology

2004· article· en· W2124082515 sur OpenAlexaboutno aff
Timothy R. Rebbeck, Christine B. Ambrosone, Douglas A. Bell, Stephen J. Chanock, Richard B. Hayes, Fred F. Kadlubar, Duncan C. Thomas

Notice bibliographique

RevueCancer Epidemiology Biomarkers & Prevention · 2004
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenetic Associations and Epidemiology
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésSingle-nucleotide polymorphismHaplotypeGenetic associationGeneticsTag SNPGenetic epidemiologyBiologySNPHaplotype estimationEpidemiologyGenome-wide association studyComputational biologyBioinformaticsMedicineGeneGenotypePathology

Résumé

récupéré en direct d'OpenAlex

The ongoing discovery of single nucleotide polymorphisms (SNPs) and characterization of haplotypes in human populations is having a fundamental impact on molecular epidemiology. While likely common polymorphic variants interact with exposures to cause human cancer, the ability to evaluate the role of SNPs in human disease is limited by available methodologies. Numerous association studies of SNPs or haplotypes have been published to elucidate the etiology of cancer, but there has been inconsistency in the ability to replicate results. Therefore, a major goal of the field is to develop the analytical tools needed to examine the explosion of genetic information available for relating genetic variants to well-defined epidemiological end points. Inconsistent methodology and results from association studies have deflected attention away from the need to establish sound methodologies for both execution and interpretation of association study data. Therefore, we must optimize epidemiological, statistical, and laboratory approaches to achieve credible outcomes in association studies.The goal of the AACR-sponsored conference “SNPs, Haplotypes, and Cancer: Applications in Molecular Epidemiology” was to address methodological developments for epidemiological studies investigating complex interrelationships of SNPs, haplotypes, and environmental factors with cancer. As summarized below, the content of this meeting included discussions related to the following:Approximately ten million SNPs exist in the human genome, with an estimated two common missense variants per gene (e.g., 1). At least 5 million SNPs have already been reported in public databases (2). However, the ability to apply these SNPs in association studies is limited by problems validating a SNP's identity, characterizing its occurrence in relevant populations, and understanding its function. Likely, only a small subset (perhaps 50,000–250,000) of the total number of SNPs in the human genome will actually confer small to moderate effects on phenotypes that are causally related to disease risk (3).At the meeting, Stephen Sherry (National Center for Biotechnology Information) suggested several classes of variants to consider, including SNPs, deletion/insertion polymorphisms (DIP), simple tandem repeat (STR) polymorphisms, named polymorphisms (e.g., Alu/− dimorphisms), and multinucleotide polymorphisms (MNP). Of these, ∼3 million SNPs are estimated to be within or 2 kb upstream or downstream from a gene. Sherry reported a “snapshot” of gene-centric SNPs in the dbSNP database (http://www.ncbi.nlm.nih.gov/SNP) as of September 2003. The distribution of these variants was 63% intronic, 11% untranslated region, 1% nonsynonymous, 1% synonymous, 24% locus region, <1% splice site, and <1% unknown coding variant. Recent surveys of human genetic diversity have estimated that there are about 100,000–300,000 SNPs in protein coding sequences (cSNPs) of the entire human genome (1). Coding SNPs are of particular interest because some of them, termed nonsynonymous SNPs (nsSNPs) or missense variants, introduce amino acid changes into their encoded proteins. nsSNPs constitute about 1% of all SNPs. The rarity of nsSNPs may be a consequence of selective pressures. However, a significant fraction of functionally important molecular diversity in the human population likely is attributable to the effects on protein function caused by nsSNPs or alterations in the regulation of genes (known as rSNPs). For example, the kinetic parameters of enzymes, the DNA binding properties of proteins that regulate transcription, the signal transduction activities of transmembrane receptors, and the architectural roles of structural proteins are all susceptible to perturbation by nsSNPs and their associated amino acid polymorphisms. Similarly, John Potter (Fred Hutchinson Cancer Research Center) argued that perturbations in the regulation of key elements of a pathway could influence the risk for cancer outcomes either directly or through interacting pathways. Tom Hudson (McGill University) added that a major challenge of future studies will be to identify and characterize regulatory variants. The researcher will have to ask whether the regulatory SNP itself or the haplotype in which it may be imbedded is really the functional unit of interest.The identification of biologically meaningful, disease-causing variants from among this large amount of genomic variability is a key challenge for association studies. Joel Hirschhorn (Massachusetts Institute of Technology) invoked the common disease, common gene hypothesis (4, 5) using methods that rely on knowledge of candidate genes or methods that rely on linkage disequilibrium (LD). Other publications have espoused the importance of rare variants as determinants for common diseases (4). Candidate gene approaches have the advantage of maximizing inferences about biological plausibility and disease causality. However, candidate approaches are limited by the amount of information that is available about the function of the gene in a specific disease process. Alternatively, genome-wide approaches have the advantage of scanning the entire genome for associations without having to rely on choosing a priori candidates. With advances in high-throughput technology and genome-wide association methods, these approaches will be more tractable than in the past. Stephen Chanock [National Cancer Institute (NCI)] and David Hunter (Harvard University) stated that the candidate gene approach remains viable despite some limitations. The expectation of genome-wide approaches is still several years away because of formidable issues of cost and availability of genotyping platforms and analytical programs. Hunter, in a talk entitled “Death, Taxes, and Candidate Genes,” further stated that regardless of the initial approach, research must ultimately result in candidate gene studies to identify biologically meaningful causal associations involving specific genes.A major challenge in candidate gene studies is to choose appropriate candidate genes usually based on sound and plausible biologically driven hypotheses. Chanock and Stacey Gabriel (Massachusetts Institute of Technology) stressed choosing markers based on (a) strong prior information about biological pathways or linkage data; (b) functional correlates for a SNP or haplotype, including pathway or the use of evolution-based approaches to identify related genes based on sequence homology or gene family; and (c) SNP haplotype studies that start with a “simple” haplotypes (often including known nsSNPs or rSNPs), which can be expanded to increase the density of SNPs across the haplotype. Regardless of the approach for choosing markers, validation of associations in both comparable and different genetic backgrounds will be required. At the same time, the working hypothesis will likely become increasingly complex as knowledge of interrelated pathways is considered to account for relevant biological interactions.William Evans (St. Jude's Children's Hospital), Gareth Morgan (University of Leeds), and Richard Weinshilboum (Mayo Clinic) presented the pharmacogenetics and pharmacogenomics paradigm for studies of candidate gene and SNP identification, gene discovery, and genotype-environment interaction. Pharmacogenetics and pharmacogenomics are excellent paradigms for studies that extend beyond etiology to studies of treatment response, gene expression changes, survival, side effects or toxicities relating to specific agents, timing of later events, and dosing. With respect to candidate genes, Weinshilboum stated that the paradigm for functional gene discovery in pharmacogenetics began by using the distribution of phenotypic traits to infer genetic effects. More recently, it has been possible to relate functionally significant DNA sequence variation to clinically important variability. Both of these approaches are complementary and should be both done to understand the functional significance of genes and SNPs. Evans also stated that gene expression profiling can be valuable to identify and characterize candidate genes (e.g., for treatment response). Evans also presented examples in which genetic profiles differed by exposure (i.e., where combinations of drug treatments did not evoke the same expression profile as each treatment individually). Therefore, expression profile approaches may be useful for identification of novel genes, characterizing function, novel disease classifications, and studying genotype-environment interactions.Gabriel and Eric Lai (GlaxoSmithKline) noted that causal (candidate) variants need not be studied directly but that gene discovery studies can be accomplished using a strategy that relies on LD between genetic variants. This represents the underlying premise behind whole genome SNP scans. The whole genome association approach can identify new candidate genes or regions. Millions of SNPs are available from the ATLAS project (http://www.confirmant.com/indexns4.html), the SNP Consortium (http://snp.cshl.org), and through public databases such as dbSNP (http://www.ncbi.nlm.nih.gov/SNP). Over >150 million SNPs are expected to be analyzed. These data can be used to define haplotype block structures across the genome and thus facilitate selection of SNPs for whole genome analysis.Lai and Gabriel outlined approaches for undertaking genome-wide or haplotype-based studies. First, appropriate epidemiological study designs and adequate statistical power are essential. The number of samples and the number of SNPs required for these studies may be substantially higher than typical studies of the past. Second, results of genome-wide SNP association studies might not be easily replicated in subsequent studies but could still identify causative regions of the genome. Similarly, large-scale genome-wide scans may find surrogate markers that will distinguish cases from controls but may not identify causative SNPs. Therefore, replication of associations is crucial to lead to valid and causative associations. Third, LD blocks exist throughout the genome, but these blocks are of varying length and appear to vary according to differences in population genetics. For example, Stephen O'Brien (NCI) stated that LD block size tends to be shorter in individuals of African ancestry and longer in Caucasians. O'Brien also reported that ∼400,000 conserved sequence blocks exist in the human genome, which have been established by evolutionary constraints, specifically across species. Within these LD blocks, there is strong allelic association and limited haplotype diversity. Where haplotype diversity exists, particularly informative SNPs that best characterize a haplotype (tagSNPs) can be used to limit the amount of laboratory and analytical work in haplotype-based studies. Fourth, use of haplotype block information has been proposed to increase power 15–50% compared with a SNP-based analysis (6, 8). However, complete (and resource-intensive) studies of SNPs in a region are required to achieve sufficient statistical power. The alternative of studying incomplete sets of SNPs in a genomic region may result in less power but still identify causative loci. In this regard, several questions remain: What level of genomic coverage (i.e., how many SNPs) is required to achieve an adequate result? Are tagSNP approaches adequate? How well do haplotype blocks need to be characterized and in what populations before tagSNPs can be reliably used? As an intermediate approach, Gabriel suggested a survey approach of candidate genes that encompassed 100 kb surrounding 200 candidate genes with 1 SNP (each with >5% minor allele frequency)/5 kb.After introductory remarks justifying a common SNP haplotype-based association approach toward assessing risks associated with genomic variation in candidate genes, Daniel Stram (University of Southern California, Los Angeles, CA) introduced a formal statistical measure (Rh2) of the predictability of haplotypes based on genotypes and described the use of this criterion for optimally picking haplotype tagging SNPs to be genotyped in large case-control studies. Stram described two approaches for estimating haplotype-specific risks in case-control studies (7, 8). The first is to compare haplotype frequencies in cases and controls separately. Analysis must allow for in of haplotype frequencies and for Second, methods exist in which an expected haplotype is as a (e.g., using haplotype in as it to the haplotype. the alternative in of effects can be and the of related to the formal measure of haplotype (i.e., Stram presented data that in using the methods are small to an adequate of haplotype tagging SNPs is However, increase as haplotype tagging SNPs are included in The of these and approaches to study haplotype data in samples of individuals (e.g., in case-control or will facilitate the of haplotypes in association association approaches should be in the years to but formidable in and Candidate gene approaches will to causal of specific genes in regions by genome-wide or the from genome-wide studies are to candidate gene studies and Similarly, knowledge of the functional significance of SNPs is key to understanding the biological of an epidemiological function can be in for candidate gene studies or the identification of novel genes from genome-wide association (National Institute of noted that (i.e., many and (i.e., many approaches are available to molecular epidemiological studies. proposed that an genotyping approach should of SNPs to be genotyped on both DNA (e.g., using with and be to the number of Lai and Gabriel that approaches (e.g., per cost as as are but this is not for The single cost is that of cost with the number of to be The appropriate to genotyping cost is to total and by the number of genotypes major for laboratory approaches in association studies is and noted that can result in (e.g., toward the hypothesis for laboratory problems DNA or of of and suggested to address these Chanock and Gabriel suggested that may be by markers or study samples to and identify For example, the CA) of markers, which can be on to or Similarly, of to of DNA can to identify which may or of and of can identify from expected and should be included in genotyping For example, the that controls be included in and controls should be the genotyping to the for genotyping of appropriate DNA methods should be to appropriate DNA and a approach for laboratory data is the appropriate use of approaches to data to in data. The use of a is for data particularly these can be used to increase the of association studies. that for and SNPs, genotypes are required without but 2 genotypes are required for cases and controls using replicate which may also be required to account for and are required to DNA from each in may be and because DNA is required. For example, estimated that it to and 100 can the amount of DNA from each is not can the allele is from the approach is to reported that is than this may vary from laboratory to As information in the and fraction Therefore, size must be the number of to be may not as because many more may be required to Lai noted that DNA may be a useful for can be limited because are of or from the is In can only be all of the samples are all in This a study before approaches can be these DNA as an of the approaches that may be required to large-scale association large amount of genomic information is available on public databases that can be of to undertaking association studies. For example, Sherry presented the information available dbSNP (http://www.ncbi.nlm.nih.gov/SNP). (NCI) and Sherry reported that genome information from data and genome the for database is For example, SNPs reported in of coding regions variants that are to and are not SNPs Similarly, of SNPs are not and may not In SNP are not and are based on information SNP frequencies can vary substantially by SNP frequencies may not be useful frequencies are not reported by a public SNP databases SNPs. In many SNP are not that are results by genotyping methods, including of including laboratory methods and in are required to a SNP before can be reliably also reported that databases such as can be used to what genes have been and functional The database also and data tools to evolutionary and an for data analysis by using their data public the for in on public studied candidate and that the and validation This to and represents a higher density than of SNPs by analysis of not reported in These data that the density of unknown SNPs is higher than and many are not reported in public this is the will be to unknown SNPs within the of interest or are within the used and are not considered in This may lead to to or result in This of variability in and execution may genotyping that to inconsistency of results among association must be to that may result in and in association studies. approaches for data and data methods should be This can for laboratory to or replicate genotyping and key in association studies is the ability to replicate association study of association studies is required not only to identify biologically plausible causative associations but also to that a candidate gene has meaningful Hirschhorn that associations are not This of replication can be by (e.g., (e.g., studies that are to identify the or population differences (e.g., the associations are different because of differences in genetic the of in association what level of can we have in associations reported to address this Hirschhorn reported on that included associations and studies (i.e., by the initial initial associations not but an of replicated associations expected the This replication is not to because have to that studies not reported than the of reported Hirschhorn also that it was that these to LD or population or could also of but this was to be a significant of study The first also to be for reported to the which that the initial the of associations studied for a consequence of this is that replication studies may because the replication effects may be than suggested by the initial these these data that many associations are and may causative effects on achieve association must factors that influence the and interpretation of these studies. must be established for on associations do or do not exist based in on the issues outlined Numerous stated that the etiology of human disease is complex and the diseases are Therefore, association study methods need to address this (NCI) and Hunter noted that effects may different of effects and the in which a is likely to be The power and of association studies may be genetic effects are studied in by exposure particularly in genotype-environment studies. Other approaches that can to address studying of cases by or to using methods that allow of or by advantage of intermediate end (e.g., and to than studies to and with have including or of studies. Potter stated that large studies have case-control but these must have adequate risk and power to evaluate disease In this proposed to develop which or more individuals with complete data and to be paradigms for undertaking large studies large case-control studies (e.g., of or case-control studies (e.g., the Consortium for and large studies (e.g., and designs (e.g., these data genotyping technology and statistical methods have to but the population required to address relevant research questions of interest have not as the to large-scale is from large studies have the to small and clinically results that may not be a is noted that large studies have the advantage of to small effects and may the of studies may be more for gene discovery and replication than a large studies. studies of rare may be particularly subset are to be However, studies study is including of genotyping and data across and by and study and data differences across The for of data are also there are or methodological problems with the studies. Therefore, several issues need to be before large-scale studies can be to achieve replication among association studies may be to of the study Hunter the that replication that studies be comparable (e.g., in of or which is not usually the Numerous the of genetic of populations related to or (University of noted that are the least and thus to major problems in and University) reported on population genetic that among particularly African genomic to population The that of the of this population is that by (i.e., population may lead to inferences from association studies. individuals of presented data that and are not important of particularly in populations of epidemiological methods are used because differences in disease risks are not large to confer significant This is with a research by several For example, and argued that study may be important than population in to association analytical approaches exist to either problems by population genetic or that use this in gene suggested that population can be by using methods, or by markers as In the approach, association is to population there should be associations across markers across the genome. approaches exist that can this The approach a of individuals are their from different populations or This approach information from markers to infer their ancestry and about population further this population information to the association that is The approach the distribution of association for the markers to the of association for the caused by population advantage of population can also association studies. O'Brien reported on the by LD approach, which the number of SNPs that are needed for a whole genome association because LD tends to extend a in This approach differences in disease risk by and studies markers that also by can be studied to the of (e.g., of African using a and that analysis with data to identify genomic regions that appear to between cases and Similarly, (University of Southern and haplotype approaches to inferences about the evolutionary of SNP and haplotype These approaches can how have in specific populations and may identify study and analysis issues have been it remains to results of many association studies to that a SNP is is causally associated with approaches have been suggested to evaluate the of a particular this by that possible can be the an association is the association is important association can or can be proposed a to in this The on prior and size of the prior is the is and an association is more likely to be prior is size only the of a associations are more likely to be the prior of an is In small studies may be more likely to a The can be to with the ability the of of Second, the must choose a prior For example, the can ask what is the that there is a causal (e.g., risk between SNP and disease before data This prior can be based on genomic (i.e., of and and epidemiological data (i.e., and information about the of and with diseases of likely Both and Chanock that future candidate gene studies may prior as association studies to candidate genes about which less is known a Third, the must choose a clinically or meaningful size (e.g., risk or and the for each SNP and a of The can that the result is the some Similarly, the can a to in there is As on prior and suggested the prior of association for a single SNP on the of haplotype blocks, with prior in the of 1 in to 1 in For candidate gene associations are haplotype blocks exist per and there are candidate variants and of these are the prior for association is 1 in 100 to 1 in candidate gene association studies are to have a substantially higher prior of an than SNP association study and adequate statistical power are crucial to meaningful association results. of and prior or of and population should be in association studies. of and association to the but such database to this to the of associations by several including Tom Cancer and These should be into the and interpretation of all future association studies.The ability to these and will whether association studies involving SNPs or haplotypes will meaningful information about disease etiology or and thus whether this information can be further into cancer or studies that elucidate disease

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,149
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0030,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0010,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,028
Tête enseignante GPT0,349
Écart entre enseignants0,321 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations58
Publié2004
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueCancer Epidemiology Biomarkers & PreventionMême sujetGenetic Associations and EpidemiologyTravaux en français237 207