Rapid assessment of phytoplankton assemblages using Next Generation Sequencing – Barcode of Life database: a widely applicable toolkit to monitor biodiversity and harmful algal blooms (HABs)
Bibliographic record
Abstract
Abstract Harmful algal blooms have important implications for the health, functioning and services of aquatic ecosystems. Our ability to detect and monitor these events is often challenged by the lack of rapid and cost-effective methods to identify bloom-forming organisms and their potential for toxin production, Here, we developed and applied a combination of DNA barcoding and Next Generation Sequencing (NGS) for the rapid assessment of phytoplankton community composition with focus on two important indicators of ecosystem health: toxigenic bloom-forming cyanobacteria and impaired planktonic biodiversity. To develop this molecular toolset for identification of cyanobacterial and algal species present in HABs (Harmful Algal Blooms), hereafter called HAB-ID, we optimized NGS protocols, applied a newly developed bioinformatics pipeline and constructed a BOLD (Barcode of Life Data System) 16S reference database from cultures of 203 cyanobacterial and algal strains representing 101 species with particular focus on bloom and toxin producing taxa. Using the new reference database of 16S rDNA sequences and constructed mock communities of mixed strains for protocol validation we developed new NGS primer set which can recover 16S from both cyanobacteria and eukaryotic algal chloroplasts. We also developed DNA extraction protocols for cultured algal strains and environmental samples, which match commercial kit performance and offer a cost-efficient solution for large scale ecological assessments of harmful blooms while giving benefits of reproducibility and increased accessibility. Our bioinformatics pipeline was designed to handle low taxonomic resolution for problematic genera of cyanobacteria such as the Anabaena-Aphanizomenon-Dolichospermum species complex, two clusters of Anabaena (I and II), Planktothrix and Microcystis . This newly developed HAB-ID toolset was further validated by applying it to assess cyanobacterial and algal composition in field samples from waterbodies with recurrent HABs events.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".