More discussion of minimalist species descriptions and clarifying some misconceptions contained in Meier et al. 2021
Bibliographic record
Abstract
This is a response to a preprint version of “A re-analysis of the data in Sharkey et al.’s (2021) minimalist revision reveals that BINs do not deserve names, but BOLD Systems needs a stronger commitment to open science”, https://www.biorxiv.org/content/10.1101/2021.04.28.441626v2. Meier et al. strongly criticized Sharkey et al.’s publication in which 403 new species were deliberately minimally described, based primarily on COI barcode sequence data. Here we respond to these criticisms. The following points are made: 1) Sharkey et al. did not equate BINs with species, as demonstrated in several examples in which multiple species were found to be in single BINs. 2) We reiterate that BINs were used as a preliminary sorting tool, just as preliminary morphological identification commonly sorts specimens based on color and size into unit trays; despite BINs and species concepts matching well over 90% of species, this matching does not equate to equality. 3) Consensus barcodes were used only to provide a diagnosis to conform to the rules of the International Code of Zoological Nomenclature just as consensus morphological diagnoses are. The barcode of a holotype is definitive and simply part of its cellular morphology. 4) Minimalist revisions will facilitate and accelerate future taxonomic research, not hinder it. 5) We refute the claim that the BOLD sequences of Plesiocoelus vanachterbergi are pseudogenes and demonstrate that they simply represent a frameshift mutation. 6) We reassert our observation that morphological evidence alone is insufficient to recognize species within species-rich higher taxa and that its usefulness lies in character states that are congruent with molecular data. 7) We show that in the cases in which COI barcodes code for the same amino acids in different putative species, data from morphology, host specificity, and other ecological traits reaffirm their utility as indicators of genetically distinct lineages.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".