Pursuit for est microsatellites in a tetraploid model from de novo transcriptome sequencing
Bibliographic record
Abstract
Available scientific literature reports very few microsatellite markers derived from tetraploid genomes using de novo transcriptome sequencing, mostly because their gain usually represents a major computational challenge due to complicated combinatorics during assembly of sequence reads. Here we present a novel approach for mining polymorphic microsatellite loci from transcriptome data in a tetraploid species with no reference genome available. Pairs of 114 bp long de novo sequenced transcriptome reads of Centaurium erythraea were merged into short contigs of 170-200 bp each. High accuracy assembly of the pairs of reads was accomplished by a minimum of 14 bp overlap. Sequential bioinformatics operations involved fully free and open-source software and were performed using an average personal computer. Out of the 13 150 candidate contigs harboring SSR motifs obtained in a final output, we randomly chose 16 putative markers for which we designed primers. We tested the effectiveness of the established bioinformatics approach by amplifying them in eight different taxa within the genus Centaurium having various ploidy levels (diploids, tetraploids and hexaploids). Nine markers displayed polymorphism and/or transferability among studied taxa. They provided 54 alleles in total, ranging from 2 to 14 alleles per locus. The highest number of alleles was observed in C. erythraea, C. littorale and a hybridogenic taxon C. pannonicum. The developed markers are qualified to be used in genetic population studies on declining natural populations of Centaurium species, thus providing valuable information to evolutionary and conservation biologists. The developed cost-effective methodology provides abundant de novo assembled short contigs and holds great promise to mine numerous additional EST-SSR-containing markers for possible use in genetics population studies of tetraploid taxa within the genus Centaurium.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".