Transcriptome characterization via 454 pyrosequencing of the annelid Pristina leidyi, an emerging model for studying the evolution of regeneration
Bibliographic record
Abstract
BACKGROUND: The naid annelids contain a number of species that vary in their ability to regenerate lost body parts, making them excellent candidates for evolution of regeneration studies. However, scant sequence data exists to facilitate such studies. We constructed a cDNA library from the naid Pristina leidyi, a species that is highly regenerative and also reproduces asexually by fission, using material from a range of regeneration and fission stages for our library. We then sequenced the transcriptome of P. leidyi using 454 technology. RESULTS: 454 sequencing produced 1,550,174 reads with an average read length of 376 nucleotides. Assembly of 454 sequence reads resulted in 64,522 isogroups and 46,679 singletons for a total of 111,201 unigenes in this transcriptome. We estimate that over 95% of the transcripts in our library are present in our transcriptome. 17.7% of isogroups had significant BLAST hits to the UniProt database and these include putative homologs of a number of genes relevant to regeneration research. Although many sequences are incomplete, the mean sequence length of transcripts (isotigs) is 707 nucleotides. Thus, many sequences are large enough to be immediately useful for downstream applications such as gene expression analyses. Using in situ hybridization, we show that two Wnt/β-catenin pathway genes (homologs of frizzled and β-catenin) present in our transcriptome are expressed in the regeneration blastema of P. leidyi, demonstrating the usefulness of this resource for regeneration research. CONCLUSIONS: 454 sequencing is a rapid and efficient approach for identifying large numbers of genes in an organism that lacks a sequenced genome. This transcriptome dataset will be a valuable resource for molecular analyses of regeneration in P. leidyi and will serve as a starting point for comparisons to non-regenerating naids. It also contributes significantly to the still limited genomic resources available for annelids and lophotrochozoans more generally.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".