Bibliographic record
Abstract
A dataset containing 391134 species occurrences available in GBIF matching the query: { "and" : [ { "or" : [ "BasisOfRecord is Living Specimen", "BasisOfRecord is Literature Occurrence", "BasisOfRecord is Material sample", "BasisOfRecord is Human Observation", "BasisOfRecord is Machine Observation", "BasisOfRecord is Observation" ] }, { "or" : [ "Continent is Africa", "Continent is Antarctica", "Continent is Asia", "Continent is Oceania", "Continent is Europe", "Continent is North America", "Continent is South America" ] }, { "or" : [ "Country is Botswana", "Country is Ethiopia", "Country is Uganda", "Country is Morocco", "Country is Tanzania, United Republic of", "Country is Mauritania", "Country is Kenya", "Country is Namibia", "Country is Benin", "Country is South Africa", "Country is Antarctica", "Country is South Georgia and the South Sandwich Islands", "Country is Cambodia", "Country is Russian Federation", "Country is India", "Country is Israel", "Country is Sri Lanka", "Country is New Zealand", "Country is Australia", "Country is Netherlands", "Country is Sweden", "Country is Spain", "Country is Denmark", "Country is Ireland", "Country is Belgium", "Country is Portugal", "Country is France", "Country is United Kingdom of Great Britain and Northern Ireland", "Country is Iceland", "Country is Greece", "Country is Switzerland", "Country is Poland", "Country is Svalbard and Jan Mayen", "Country is Romania", "Country is Andorra", "Country is Finland", "Country is Norway", "Country is Germany", "Country is Bulgaria", "Country is Georgia", "Country is Costa Rica", "Country is Canada", "Country is United States of America", "Country is Greenland", "Country is Estonia", "Country is Hungary", "Country is Brazil", "Country is Ecuador", "Country is Uruguay", "Country is Chile", "Country is Colombia", "Country is Argentina" ] }, "OccurrenceStatus is Present", "TaxonKey is Carnivora" ] } The dataset includes 391134 records from 198 constituent datasets; see https://api.gbif.org/v1/occurrence/download/0182031-200613084148143/datasets/export for details. Data from some individual datasets included in this download may be licensed under less restrictive terms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.009 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.288 | 0.357 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".