Bibliographic record
Abstract
This dataset observes how plant abundance and location affect insect abundance. This experiment was completed by myself, Adamo, Ava, Ashley and Katherine. The project ran on two separate days from approximately 3:00 to 5:00 pm, at York University in Toronto, Canada. The two days were October 13th and October 20th. The data was collected in two different locations. On day 1, 30 quadrats were placed in the Danby woodlot, and 30 quadrats were placed in the Danby grassland adjacent to the woodlot. The first day had clear weather, no rain and a temperature of 19⁰C. The second day, a week after day 1, the same process was repeated providing a total of 60 quadrats of data per location. On day 2, there was light rain during data collection, and a temperature of 18⁰C. To place the quadrat, a random number generator was used to determine how many steps should be taken. The steps were measured by the same person each time to ensure consistency. The plant abundance was rated on a scale excluding grass and dead plants, and then the insects were counted by all members of the group. In the grassland, insect abundance ranged from 0- 18 insects counted per quadrat. The most common number of bugs to be observed was around 5-7. In the woodlot, insect abundance ranged from 0-30 insects per quadrat. The most common number of bugs to be observed was 4.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.009 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".