Calibrating the taxonomy of a megadiverse insect family: 3000 DNA barcodes from geometrid type specimens (Lepidoptera, Geometridae)
Bibliographic record
Abstract
It is essential that any DNA barcode reference library be based upon correctly identified specimens. The Barcode of Life Data Systems (BOLD) requires information such as images, geo-referencing, and details on the museum holding the voucher specimen for each barcode record to aid recognition of potential misidentifications. Nevertheless, there are misidentifications and incomplete identifications (e.g., to a genus or family) on BOLD, mainly for species from tropical regions. Unfortunately, experts are often unavailable to correct taxonomic assignments due to time constraints and the lack of specialists for many groups and regions. However, considerable progress could be made if barcode records were available for all type specimens. As a result of recent improvements in analytical protocols, it is now possible to recover barcode sequences from museum specimens that date to the start of taxonomic work in the 18th century. The present study discusses success in the recovery of DNA barcode sequences from 2805 type specimens of geometrid moths which represent 1965 species, corresponding to about 9% of the 23 000 described species in this family worldwide and including 1875 taxa represented by name-bearing types. Sequencing success was high (73% of specimens), even for specimens that were more than a century old. Several case studies are discussed to show the efficiency, reliability, and sustainability of this approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".