Bibliographic record
Abstract
A dataset containing 955 species occurrences available in GBIF matching the query: { "and" : [ "Country is Canada", "DatasetKey is iNaturalist Research-grade Observations", "Geometry POLYGON((-114.00339 51.03559,-114.00502 51.03568,-114.00622 51.03564,-114.00712 51.03658,-114.00768 51.03679,-114.01184 51.03697,-114.01257 51.03688,-114.01283 51.03637,-114.0118 51.03499,-114.01077 51.03435,-114.00983 51.03323,-114.00854 51.03216,-114.00553 51.03126,-114.0045 51.03091,-114.00768 51.03027,-114.00764 51.02967,-114.01146 51.02933,-114.01133 51.02293,-114.01116 51.0216,-114.01197 51.01864,-114.01129 51.018,-114.01043 51.01585,-114.01094 51.01448,-114.01021 51.01298,-114.00953 51.01581,-114.00914 51.01684,-114.00832 51.0183,-114.00682 51.01971,-114.00523 51.02152,-114.00262 51.02405,-114.00116 51.02516,-113.99867 51.02688,-113.99798 51.02765,-113.99682 51.02954,-113.99609 51.03126,-113.99652 51.03315,-113.99687 51.03396,-113.99785 51.03469,-114.00026 51.03654,-114.00073 51.03692,-114.00287 51.03658,-114.0042 51.03572,-114.00339 51.03559))", "OccurrenceStatus is Present" ] } The dataset includes 955 records from 1 constituent datasets; see https://api.gbif.org/v1/occurrence/download/0263262-200613084148143/datasets/export for details. Data from some individual datasets included in this download may be licensed under less restrictive terms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.010 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.288 | 0.449 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".