Comparing Approaches to Specimen Identification using Neotropical Freshwater Fishes in the Barra del Colorado Wildlife Refuge, Costa Rica
Bibliographic record
Abstract
Abstract As global biodiversity declines continue, conservation efforts are increasingly important in megadiverse areas such as the Neotropics where biodiversity is especially imperiled. The accurate identification of specimens is critical to successful conservation plans. However, in groups such as freshwater fishes, different identification methodologies have documented challenges. Using a biodiversity survey of fishes from the Barra del Colorado Wildlife Refuge in northeastern Costa Rica, we compared: (1) morphological identifications in the field, (2) morphological identifications in the lab by experts, (3) DNA barcode-based identifications, and (4) identifications based on an integrative approach. Our results suggest that both barcode-based identifications and field morphological identifications provided fewer correct species identifications than lab identifications performed by experts using morphology. We attribute shortfalls of DNA barcoding in this case to the misidentification of reference material, the use of outdated taxonomy for references sequences, and the non-uniform representation of groups in public databases across taxa. We suggest the use of an integrative approach to identify freshwater fishes in Costa Rica and other megadiverse areas of the Neotropics where similar issues with public barcode reference libraries exist. We also recommend the creation of regional curated barcode reference libraries to aid in the identification of traditionally difficult to identify species/specimens. We also provide the most up to date species list for the ichthyofauna of the Barra del Colorado Wildlife Refuge identifying 51 species from 42 genera, 21 families, and 17 orders. Generating accurate species lists for protected areas and areas of importance will provide conservation practitioners with effective tools for tracking diversity changes over time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".