Spatial variation in bicycling risk based on crowdsourced safety data
Bibliographic record
Abstract
Bicycling‐related injury data are difficult to obtain from official reports, which capture only about 20% of crashes and often lack coordinates, injury outcomes, and narratives needed for understanding where and why incidents occurred. Crowdsourced data on bicycling safety provides new opportunities for the study of bicycling injury and risk. Our goal was to quantify factors that influence the spatial variation in unsafe bicycling across a city, based on self‐reports of bicycling incidents. To meet this goal, we leveraged BikeMaps.org , a global tool for reporting bicycling safety incidents, drawing on data from Metro Vancouver. We summarized incident conditions that led to injury, developed a model to identify predictors of injury using random forest regression, and mapped bicycling incident hot spots. Our results demonstrate that injuries from bicycling incidents are associated with older and younger bicyclists, downhill slopes, parked cars, recreation and weekend rides, falls, and single bicycle incidents with infrastructure, roads, and railroads. The broad range of incidents reported to BikeMaps.org allows us to add evidence that falls and single bicycle collisions are major causes of injury. Also, we demonstrate the value of attributing safety hot spots with contextual details to identify infrastructure interventions that can reduce injury for bicyclists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.027 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".