What does crowdsourced data tell us about bicycling injury? A case study in a mid-sized Canadian city
Bibliographic record
Abstract
With only ∼20 % of bicycling crashes captured in official databases, studies on bicycling safety can be limited. New datasets on bicycling incidents are available via crowdsourcing applications, with opportunity for analyses that characterize reporting patterns. Our goal was to characterize patterns of injury in crowdsourced bicycle incident reports from BikeMaps.org. We extracted 281 incidents reported on the BikeMaps.org global mapping platform and analyzed 21 explanatory variables representing personal, trip, route, and crash characteristics. We used a balanced random forest classifier to classify three outcomes: (i) collisions resulting in injury requiring medical treatment; (ii) collisions resulting in injury but the bicyclist did not seek medical treatment; and (iii) collisions that did not result in injury. Results indicate the ranked importance and direction of relationship for explanatory variables. By knowing conditions that are most associated with injury we can target interventions to reduce future risk. The most important reporting pattern overall was the type of object the bicyclist collided with. Increased probability of injury requiring medical treatment was associated with collisions with animals, train tracks, transient hazards, and left-turning motor vehicles. Falls, right hooks, and doorings were associated with incidents where the bicyclist was injured but did not seek medical treatment, and conflicts with pedestrians and passing motor vehicles were associated with minor collisions with no injuries. In Victoria, British Columbia, Canada, bicycling safety would be improved by additional infrastructure to support safe left turns and around train tracks. Our findings support previous research using hospital admissions data that demonstrate how non-motor vehicle crashes can lead to bicyclist injury and that route characteristics and conditions are factors in bicycling collisions. Crowdsourced data have potential to fill gaps in official data such as insurance, police, and hospital reports.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".