Data set of 1,275 images capturing interactions between flies and blooming flowers
Bibliographic record
Abstract
The dataset presented is a collection of 1,275 images capturing interactions between flies and blooming flowers. The images were sourced from internet repositories through searches conducted between August 2016 and August 2020, using the Google Chrome v. 33.x web browser. Internet searches focused on Google Images and three major social media platforms: Flickr, Instagram, and agefotostock. Photographs were taken by photographers worldwide and uploaded to these platforms, forming the basis of the dataset. The data encompasses various taxonomic and ecological information for both the flies and the flowers depicted in the images. For each image, detailed taxonomic information was recorded for the flies, including their suborders (Nematocera and Brachycera, grouped as Higher/Lesser), Family (wherever possible, distinguishing between Syrphidae and non-Syrphidae), and Genus and species (if available). Further characterization of the flies included recording their sex (identified based on morphology), feeding status (identified by visible proboscis extension into the flower), and the presence or absence of pollen particles on their bodies. Similarly, for each image, taxonomic information was collected for the flowers, including their Family and Genus and species (when identifiable). Flowers were categorized by petal color, which was grouped into four main categories based on the visible spectrum wavelength: purple to blue (ranging from 380-520 nm), green to yellow (ranging from 520-590 nm), orange to red (ranging from 590-740 nm), and white. Additionally, flowers were classified based on their shape, with four main categories: elongate cluster, round cluster, composite-shaped, or simple-shaped. To complement the taxonomic and morphological information, the dataset includes additional data for each image, such as web links to the original sources, geographic locations where the images were captured, and the date of image acquisition. This dataset offers a valuable resource for studying fly-flower interactions on a broad scale, using photographs contributed by photographers from around the world.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.016 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".