MétaCan
Menu
Back to cohort
Record W2089096042 · doi:10.1002/humu.10143

Proposal for an allele nomenclature system based on the evolutionary divergence of haplotypes

2002· article· en· W2089096042 on OpenAlexfundno aff
Daniel W. Nebert

Bibliographic record

VenueHuman Mutation · 2002
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Associations and Epidemiology
Canadian institutionsnot available
FundersNational Institute of Environmental Health SciencesInternational Union of Biochemistry and Molecular Biology
KeywordsHaplotypeBiologyGeneticsHomo sapiensAlleleEvolutionary biologyPopulationGene

Abstract

fetched live from OpenAlex

The classical view of what constitutes an "allele" has been challenged by recent findings of a great deal of human genetic variability, i.e., we can expect, on average, one variant site every 100-250 bases of our haploid genome. The haplotype is defined as "the patterns of co-occurrence of variant sites on the same chromosome" (and therefore within each particular gene). Sufficient evidence exists for the divergence of haplotypes during evolution of Homo sapiens sapiens, and the total number of haplotypes per gene will reflect the amount of time any particular ethnic group has existed on the planet, e.g., greatest in Africans, fewer in East Asians, and still fewer in Caucasians. If the average gene spans 30 kb, we can expect approximately 170 polymorphic variant sites per gene in the world population. We do not see 2(170) haplotypes, however; we might find only 10 to 200 haplotypes (depending on the gene's size and degree of conservation of the gene product). This finite number allows for a reasonable haplotype nomenclature system for each gene, based on evolutionary divergence. For polymorphic variants (i.e., frequency > or = 0.01), I propose using Arabic numerals for the major clades (e.g., *1, *2, em leader *20, *21), capital letters for sublineages (e.g., *2A, *2B, *2C), and Arabic numerals for sub-sublineages (e.g., *22G12, *22G13); additional subcategories may be added, in an alternating number/letter/number/letter sequence, depending on the complexity of present-day haplotypes of a particular gene. Web sites with a web master and external advisory committee should be set up for each gene superfamily, family, or individual gene (depending on complexity), and an international haplotype nomenclature committee, perhaps comprised of several dozen of these web masters, should oversee haplotype nomenclature for the entire human genome. The higher heterozygosity and multiallelic nature makes haplotypes more informative than biallelic SNPs. Ultimately, our knowledge of haplotype patterns, rather than single variant sites, of perhaps several hundred genes will likely be helpful in finding associations between genotype and any multiplex phenotype (e.g., complex diseases including cancer, and/or toxicity of pharmaceutical agents or environmental pollutants).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.505
Threshold uncertainty score0.183

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.027
GPT teacher head0.267
Teacher spread0.240 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations28
Published2002
Admission routes1
Has abstractyes

Explore more

Same venueHuman MutationSame topicGenetic Associations and EpidemiologyFrench-language works237,207