Bibliographic record
Abstract
A dataset containing 1311254 species occurrences available in GBIF matching the query: { "and" : [ "Country is Spain", { "or" : [ "PublishingCountry is Norway", "PublishingCountry is Germany", "PublishingCountry is Belgium", "PublishingCountry is Finland", "PublishingCountry is Chinese Taipei", "PublishingCountry is Russian Federation", "PublishingCountry is Portugal", "PublishingCountry is Bulgaria", "PublishingCountry is Japan", "PublishingCountry is Denmark", "PublishingCountry is Luxembourg", "PublishingCountry is France", "PublishingCountry is New Zealand", "PublishingCountry is Brazil", "PublishingCountry is Sweden", "PublishingCountry is Morocco", "PublishingCountry is Slovenia", "PublishingCountry is United Kingdom of Great Britain and Northern Ireland", "PublishingCountry is Canada", "PublishingCountry is United States of America", "PublishingCountry is Estonia", "PublishingCountry is Switzerland", "PublishingCountry is South Africa", "PublishingCountry is Korea, Republic of", "PublishingCountry is Mexico", "PublishingCountry is Italy", "PublishingCountry is Colombia", "PublishingCountry is Venezuela (Bolivarian Republic of)", "PublishingCountry is Costa Rica", "PublishingCountry is Argentina", "PublishingCountry is Austria", "PublishingCountry is Australia", "PublishingCountry is Peru", "PublishingCountry is Czechia", "PublishingCountry is Nicaragua", "PublishingCountry is Poland", "PublishingCountry is Netherlands" ] }, "HasGeospatialIssue is false" ] } The dataset includes 1311254 records from 1003 constituent datasets; see https://api.gbif.org/v1/occurrence/download/0028249-180131172636756/datasets/export for details. Data from some individual datasets included in this download may be licensed under less restrictive terms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.057 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".