Methods for DNA Barcoding Photosynthetic Protists Emphasizing the Macroalgae and Diatoms
Bibliographic record
Abstract
This chapter outlines the current practices used in our laboratory for routine DNA barcode analyses of the three major marine macroalgal groups, viz., brown (Phaeophyceae), red (Rhodophyta), and green (Chlorophyta) algae, as well as for the microscopic diatoms (Bacillariophyta). We start with an outline of current streamlined field protocols, which facilitate the collection of substantial (hundreds to thousands) specimens during short (days to weeks) field excursions. We present the current high-throughput DNA extraction protocols, which can, nonetheless, be easily modified for manual molecular laboratory use. We are advocating a two-marker approach for the DNA barcoding of protists with each major lineage having a designated primary and secondary barcode marker of which one is always the LSU D2/D3 (divergent domains D2/D3 of the nuclear ribosomal large subunit DNA). We provide a listing of the primers that we currently use in our laboratory for amplification of DNA barcode markers from the groups that we study: LSU D2/D3, which we advocate as a eukaryote-wide barcode marker to facilitate broad ecological and environmental surveys (secondary barcode marker in this capacity); COI-5P (the standard DNA barcode region of the mitochondrial cytochrome c oxidase 1 gene) as the primary barcode marker for brown and red algae; rbcL-3P (the 3' region of the plastid large subunit of ribulose-l-5-bisphosphate carboxylase/oxygenase) as the primary barcode marker for diatoms; and tufA (plastid elongation factor Tu gene) as the primary barcode marker for chlorophytan green algae. We outline our polymerase chain reaction and DNA sequencing methodologies, which have been streamlined for efficiency and to reduce unnecessary cleaning steps. The combined information should provide a helpful guide to those seeking to complete barcode research on these and related "protistan" groups (the term protist is not used in a phylogenetic context; it is simply a catch-all term for the bulk of eukaryotic diversity, i.e., all lineages excluding animals, true fungi, and plants).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.009 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".