A <scp>DNA</scp> barcode library for Germany′s mayflies, stoneflies and caddisflies (Ephemeroptera, Plecoptera and Trichoptera)
Bibliographic record
Abstract
Mayflies, stoneflies and caddisflies (Ephemeroptera, Plecoptera and Trichoptera) are prominent representatives of aquatic macroinvertebrates, commonly used as indicator organisms for water quality and ecosystem assessments. However, unambiguous morphological identification of EPT species, especially their immature life stages, is a challenging, yet fundamental task. A comprehensive DNA barcode library based upon taxonomically well-curated specimens is needed to overcome the problematic identification. Once available, this library will support the implementation of fast, cost-efficient and reliable DNA-based identifications and assessments of ecological status. This study represents a major step towards a DNA barcode reference library as it covers for two-thirds of Germany's EPT species including 2,613 individuals belonging to 363 identified species. As such, it provides coverage for 38 of 44 families (86%) and practically all major bioindicator species. DNA barcode compliant sequences (≥500 bp) were recovered from 98.74% of the analysed specimens. Whereas most species (325, i.e., 89.53%) were unambiguously assigned to a single Barcode Index Number (BIN) by its COI sequence, 38 species (18 Ephemeroptera, nine Plecoptera and 11 Trichoptera) were assigned to a total of 89 BINs. Most of these additional BINs formed nearest neighbour clusters, reflecting the discrimination of geographical subclades of a currently recognized species. BIN sharing was uncommon, involving only two species pairs of Ephemeroptera. Interestingly, both maximum pairwise and nearest neighbour distances were substantially higher for Ephemeroptera compared to Plecoptera and Trichoptera, possibly indicating older speciation events, stronger positive selection or faster rate of molecular evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".