DNA barcoding of Cuban freshwater fishes: evidence for cryptic species and taxonomic conflicts
Bibliographic record
Abstract
Despite ongoing efforts to protect species and ecosystems in Cuba, habitat degradation, overuse and introduction of alien species have posed serious challenges to native freshwater fish species. In spite of the accumulated knowledge on the systematics of this freshwater ichthyofauna, recent results suggested that we are far from having a complete picture of the Cuban freshwater fish diversity. It is estimated that 40% of freshwater Cuban fish are endemic; however, this number may be even higher. Partial sequences (652 bp) of the mitochondrial gene COI (cytochrome c oxidase subunit I) were used to barcode 126 individuals, representing 27 taxonomically recognized species in 17 genera and 10 families. Analysis was based on Kimura 2-parameter genetic distances, and for four genera a character-based analysis (population aggregation analysis) was also used. The mean conspecific, congeneric and confamiliar genetic distances were 0.6%, 9.1% and 20.2% respectively. Molecular species identification was in concordance with current taxonomical classification in 96.4% of cases, and based on the neighbour-joining trees, in all but one instance, members of a given genera clustered within the same clade. Within the genus Gambusia, genetic divergence analysis suggests that there may be at least four cryptic species. In contrast, low genetic divergence and a lack of diagnostic sites suggest that Rivulus insulaepinorum may be conspecific with Rivulus cylindraceus. Distance and character-based analysis were completely concordant, suggesting that they complement species identification. Overall, the results evidenced the usefulness of the DNA barcodes for cataloguing Cuban freshwater fish species and for identifying those groups that deserve further taxonomic attention.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".