MétaCan
Menu
Back to cohort
Record W1998365429 · doi:10.1101/gr.10.8.1071

WABA Success: A Tool for Sequence Comparison between Large Genomes

2000· article· en· W1998365429 on OpenAlexaff
David L. Baillie, Ann M. Rose

Bibliographic record

VenueGenome Research · 2000
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetics, Aging, and Longevity in Model Organisms
Canadian institutionsSimon Fraser UniversityUniversity of British Columbia
Fundersnot available
KeywordsBiologyGenomeCaenorhabditisGeneticsCaenorhabditis elegansSequence assemblyIntronWhole genome sequencingGeneEvolutionary biology

Abstract

fetched live from OpenAlex

Whole-genome sequence comparisons between bacterial sequences are one thing, but try comparing two eukaryotic genomes, each containing tens or hundreds of millions of nucleotides. And try to do it on your desktop machine in your office or at home. That is what Kent and Zahler (2000) have tried, and the results are presented in this issue of Genome Research. The use of evolutionary conservation to unveil functional information contained within genomes is not new. In the case of the nematode, comparisons of Caenorhabditis elegans to its close relative Caenorhabditis briggsae go back as far as Emmons et al. (1979). Snutch (1984) made the first C. briggsae genomic library available to the research community. These two nematodes are almost identical in morphology and development, yet their genomes have been separated for a sufficient amount of time to allow intronic and intragenic sequences to become effectively randomized, while proteinencoding sequences and cis-linked regulatory elements are conserved. C. briggsae has diverged from C. elegans in regions of unselected sequence, the middle of large introns, between genes, and at the third nucleotide of synonymous codons in genes not abundantly translated. Based on the now-outdated concept of a constant evolutionary clock, these two species were estimated to have diverged some 30–60 million years ago. These facts prompted a C. briggsae genome sequencing effort to be initiated by The Washington University Genome Sequencing Center in St. Louis, Missouri. The project was given a jumpstart with the construction of a physical map (Marra et al. 1997) and an open invitation to the research community to participate in the selection of fosmids for sequence analysis. Researchers from around the world probed microarrayed fosmid filters with their favorite gene and contacted The Genome Sequencing Center with a request to sequence the identified fosmid. Currently, approximately 10% of the genome has now been completed and is available at http://www.genome.wustl.edu/pub/ gsc1/sequence/st.louis/briggsae/finish/. Comparison of C. elegans and C. briggsae sequence has facilitated the identification of cis-linked regulatory elements (Heine and Blumenthal 1986), the cloning of genes as a result of their syntenic relationship (Kuwabara and Shah 1994), and interpretation of complex gene structure (Thacker et al. 1999). In many cases examined, adjacent genes are conserved both in position and orientation, with the occasional interruption caused by a transposable element or an apparent pseudogene. Prior to the availability of C. briggsae sequence, gene feature identification was done by abinitio prediction. Many computer programs exist that attempt to predict the exon-intron structure of genes from genomic sequence. These programs vary in their accuracy and are at their worst when asked to predict the first and last exons of genes, often failing to identify correctly exons and introns outside the actual coding element. Messenger RNA and EST-based feature detection is often confounded by large transcripts, genes that have low levels of transcription, or genes lacking poly(A)containing transcripts. Further complications adding to the problem of gene finding include the still-murky rules of sequence identification unknown for numerous important genomic elements (snRNAs, ribozymes, transcription factor binding sites, replication origins, chromatin folding and packaging signals, meiotic pairing information, etc.). Because of these issues, the initial annotation of the C. elegans genome was done interactively using GENEFINDER, a program written by Phil Green and LaDeana Hillier (unpubl.). The program, which was ahead of its time, did a remarkable job of gene prediction. Taken together with manual interpretation, the cDNA data (Kohara 1996) and related genomic sequence data from C. briggsae, gene structure predictions were made for > 19,000 C. elegans genes. Kent and Zahler (2000) have taken advantage of the availability of genomic sequence from these two closely related nematodes to test the feasibility of doing large-scale alignments between genomic DNA of different species. They have developed an algorithm for sequence comparison in which every third base (the wobble position of the codon) is ignored. The wobble-aware bulk aligner (WABA) allows the sensitive identification of conserved coding regions. They use this algorithm in a three-tiered process to identify conserved regions between eight million base pairs of C. briggsae genomic sequence and the entire 97 million base pairs of C. elegans. Their analysis was performed on a readily available 450 mHz Intel-based machine. The results they have achieved are remarkable and will provide a useful resource for all in Corresponding author. E-MAIL arose@gene.nce.ubc.ca; FAX (604) 8225348. Insight/Outlook

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.856
Threshold uncertainty score0.814

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.085
GPT teacher head0.389
Teacher spread0.303 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations15
Published2000
Admission routes1
Has abstractyes

Explore more

Same venueGenome ResearchSame topicGenetics, Aging, and Longevity in Model OrganismsFrench-language works237,207