How valid are old species lists? How archived samples can be used to update Ephemeroptera biodiversity information for northern Canada
Bibliographic record
Abstract
Abstract Broad-scale aquatic insect ecological studies are an important potential source of biodiversity information, though taxa lists may contain outdated names or be incompletely or incorrectly identified. We re-examined over 12 000 archived Ephemeroptera (mayfly) specimens from a large environmental assessment project (Mackenzie Valley pipeline study) in Yukon and the Northwest Territories, Canada (1971–1973) and compared the results to data from five recent (post-2000) collecting expeditions. Our goals were to update the species list for Ephemeroptera for Yukon and the Northwest Territories, and to evaluate the benefits of retaining and re-examining ecological samples to improve regional biodiversity information, particularly in isolated or inaccessible areas. The original pipeline study specimen labels reported 17 species in 25 genera for the combined Yukon and Northwest Territories samples, of which six species and 15 genera are still valid. Re-examination of specimens resulted in 45 species in 29 genera, with 14 and seven newly recorded species for Northwest Territories and Yukon, respectively. The recent collecting resulted in 50 species, 29 of which were different from the pipeline study, and five of which were new territorial records (Northwest Territories: four species; Yukon: one species). Re-examination of archived ecological specimens provides a cost-effective way to update regional biodiversity information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.090 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.010 | 0.009 |
| Science and technology studies | 0.005 | 0.003 |
| Scholarly communication | 0.009 | 0.006 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".