Reviewing taxonomic bias in a megadiverse country: primary biodiversity data, cultural salience, and scientific interest of South African animals
Bibliographic record
Abstract
Taxonomic bias, resulting in some taxa receiving more attention than others, has been shown to persist throughout history. Such bias in primary biodiversity data needs to be addressed because the data are vital to environmental management. This study reviews taxonomic bias in South African primary biodiversity data obtained from the Global Biodiversity Information Facility (GBIF). The focus was specifically on animal classes, and regression analysis was used to assess the influence of scientific interest and cultural salience on taxonomic bias. A higher resolution analysis of the two explanatory variables’ influence on taxonomic bias is conducted using a generalised linear model on a subset of herpetofaunal families from the focal classes. Furthermore, the potential effects of cultural salience and scientific interest on a taxon’s extinction risk are investigated. The findings show that taxonomic bias in South Africa’s primary biodiversity data has similarities with global scale taxonomic bias. Among animal classes, there is strong bias towards birds while classes such as Polychaeta and Maxillopoda are under-represented. Cultural salience has a stronger influence on taxonomic bias than scientific interest. It is, however, unclear how these explanatory variables may influence the extinction risk of taxa. We recommend that taxonomic bias can be reduced if primary biodiversity data collection has a range of targets that guide (but do not limit) accumulation of species occurrence records per habitat. Within this range, a lower target of species occurrence records accommodates species that are difficult to detect. The upper target means occurrence records for any species are less urgent but nonetheless useful and thus data collection efforts can focus on species with fewer occurrence records.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".