Documentation of an effective method for undertaking molecular genomics in a mollusc: DNA extraction, microsatellite analysis, RAD-seq, and whole genome sequencing in the Horse mussel (Modiolus modiolus)
Bibliographic record
Abstract
Molluscs, one of the most diverse and ancient phyla, are globally significant as sources of protein from fisheries and aquaculture. Many species play crucial roles in maintaining ecosystem functions and act as keystone species by engineering biogenic habitats. Despite their ecological and economic importance, molecular genomic studies on molluscs have been limited due to challenges in obtaining quality DNA samples, primarily caused by the co-purification of polyphenolic proteins and mucopolysaccharides. In this report, we document a series of effective methodologies for conducting molecular genomics in the Horse mussel (Modiolus modiolus), a bivalve mollusc known for forming complex benthic structures that enhance biodiversity and provide essential ecosystem services. We provide comprehensive laboratory protocols detailing tissue collection, DNA extraction, purification, and quantification, as well as microsatellite PCR and analysis, and the development and sequencing of restriction site-associated DNA (RAD-seq). Furthermore, we present methodologies and results for de novo whole genome sequencing using both Illumina short-read and benchtop Nanopore long-read technologies. Unmodified manufacturer instructions or standard laboratory procedures are included in full to ensure methodological completeness and facilitate reproducibility We include relevant quality metrics and results for each molecular method. Although these methods may not be universally applicable to all mollusc species, they offer a validated starting point for similar genomic studies. Our findings contribute valuable techniques and insights for advancing molluscan molecular genomic research, thereby facilitating further ecological and evolutionary studies in this critical phylum.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".