Low-cost ddRAD method of SNP discovery and genotyping applied to the periwinkle<i>Littorina saxatilis</i>
Bibliographic record
Abstract
Restriction-associated DNA sequencing methods are useful for simultaneously developing and genotyping DNA markers such as single nucleotide polymorphisms (SNPs). We describe a new inexpensive protocol for double-digest restriction-associated DNA (ddRAD) sequencing, requiring purchase of only two double-stranded adapter oligonucleotides complementary to the overhanging bases left by digestion with chosen restriction enzymes. Indexing of samples is instead achieved by incorporating two unique index sequences into the forward and reverse primer sequences, so that they can both be added with PCR. This modification enables combinatorial indexing of samples by paired-end sequencing. We tested this method by preparing individual genomic libraries from two putative parents and eight putative offspring from an experimental cross of a marine snail (Littorina saxatilis); each snail's DNA was extracted, double-digested with PstI and BglII, then ligated to adapters. More than 90% of the reads (12,175,413 paired-end reads and 24,350,826 total sequences) could be assigned to the sequenced individuals. Trimmed, paired reads from the putative parents were assembled into 3,421 contigs with an N50 of 135 bp. Reads from all individuals were aligned to the parental reference assembly, allowing discovery and validation of 1,131 variant SNP sites genotyped in all individuals, with mean coverage depth of 33.54 reads per locus. Individual genotypes at each of 1,131 loci were used in parentage analysis in COLONY 2.0.4.4 and confirmed that the putative parents were the true parents of eight sequenced offspring. This study demonstrates the utility of the new low-cost ddRAD protocol for library preparation and SNP variant discovery, and will enable flexibility in choice of restriction enzyme and decrease in startup costs of future ddRAD studies in molluscan species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".