WABA Success: A Tool for Sequence Comparison between Large Genomes
Bibliographic record
Abstract
Whole-genome sequence comparisons between bacterial sequences are one thing, but try comparing two eukaryotic genomes, each containing tens or hundreds of millions of nucleotides. And try to do it on your desktop machine in your office or at home. That is what Kent and Zahler (2000) have tried, and the results are presented in this issue of Genome Research. The use of evolutionary conservation to unveil functional information contained within genomes is not new. In the case of the nematode, comparisons of Caenorhabditis elegans to its close relative Caenorhabditis briggsae go back as far as Emmons et al. (1979). Snutch (1984) made the first C. briggsae genomic library available to the research community. These two nematodes are almost identical in morphology and development, yet their genomes have been separated for a sufficient amount of time to allow intronic and intragenic sequences to become effectively randomized, while proteinencoding sequences and cis-linked regulatory elements are conserved. C. briggsae has diverged from C. elegans in regions of unselected sequence, the middle of large introns, between genes, and at the third nucleotide of synonymous codons in genes not abundantly translated. Based on the now-outdated concept of a constant evolutionary clock, these two species were estimated to have diverged some 30–60 million years ago. These facts prompted a C. briggsae genome sequencing effort to be initiated by The Washington University Genome Sequencing Center in St. Louis, Missouri. The project was given a jumpstart with the construction of a physical map (Marra et al. 1997) and an open invitation to the research community to participate in the selection of fosmids for sequence analysis. Researchers from around the world probed microarrayed fosmid filters with their favorite gene and contacted The Genome Sequencing Center with a request to sequence the identified fosmid. Currently, approximately 10% of the genome has now been completed and is available at http://www.genome.wustl.edu/pub/ gsc1/sequence/st.louis/briggsae/finish/. Comparison of C. elegans and C. briggsae sequence has facilitated the identification of cis-linked regulatory elements (Heine and Blumenthal 1986), the cloning of genes as a result of their syntenic relationship (Kuwabara and Shah 1994), and interpretation of complex gene structure (Thacker et al. 1999). In many cases examined, adjacent genes are conserved both in position and orientation, with the occasional interruption caused by a transposable element or an apparent pseudogene. Prior to the availability of C. briggsae sequence, gene feature identification was done by abinitio prediction. Many computer programs exist that attempt to predict the exon-intron structure of genes from genomic sequence. These programs vary in their accuracy and are at their worst when asked to predict the first and last exons of genes, often failing to identify correctly exons and introns outside the actual coding element. Messenger RNA and EST-based feature detection is often confounded by large transcripts, genes that have low levels of transcription, or genes lacking poly(A)containing transcripts. Further complications adding to the problem of gene finding include the still-murky rules of sequence identification unknown for numerous important genomic elements (snRNAs, ribozymes, transcription factor binding sites, replication origins, chromatin folding and packaging signals, meiotic pairing information, etc.). Because of these issues, the initial annotation of the C. elegans genome was done interactively using GENEFINDER, a program written by Phil Green and LaDeana Hillier (unpubl.). The program, which was ahead of its time, did a remarkable job of gene prediction. Taken together with manual interpretation, the cDNA data (Kohara 1996) and related genomic sequence data from C. briggsae, gene structure predictions were made for > 19,000 C. elegans genes. Kent and Zahler (2000) have taken advantage of the availability of genomic sequence from these two closely related nematodes to test the feasibility of doing large-scale alignments between genomic DNA of different species. They have developed an algorithm for sequence comparison in which every third base (the wobble position of the codon) is ignored. The wobble-aware bulk aligner (WABA) allows the sensitive identification of conserved coding regions. They use this algorithm in a three-tiered process to identify conserved regions between eight million base pairs of C. briggsae genomic sequence and the entire 97 million base pairs of C. elegans. Their analysis was performed on a readily available 450 mHz Intel-based machine. The results they have achieved are remarkable and will provide a useful resource for all in Corresponding author. E-MAIL arose@gene.nce.ubc.ca; FAX (604) 8225348. Insight/Outlook
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".