Computer Vision Techniques for the Automated Collection of Cyclist Data
Bibliographic record
Abstract
One of the main challenges in the conduct of detailed analysis of cyclist behavior is the lack of reliable data. Collection of data through manual methods is a labor-intensive and time-consuming process. Two of the important areas of cyclist data collection are volume counts and average speed measurements. A volume count provides the basis for necessary exposure measures and conveys essential information about traffic patterns. Cyclist speed data are used for traffic control and safety studies. The application of computer vision (CV) techniques enables the collection of precise spatial and temporal measurements of road users in a resource-efficient way. This paper presents the use of a set of CV techniques for the automated collection of cyclist data. Cyclist tracks obtained from video analysis were used to perform screen line counts as well as cyclist speed measurements. The applications were demonstrated with the use of a real-world data set from a roundabout in Vancouver, British Columbia, Canada. Further analysis was conducted on the mean speed of cyclists with regard to several factors (e.g., travel path, helmet use, group size). The motivation for this research was to understand better cyclist behavior and how it varied under different conditions. Several conclusions could be drawn from the analysis of cyclist speed behavior. Group size, travel path, lane position, and helmet use were all found to affect the cyclist mean speed. Single cyclists had a slightly, but significantly, higher mean cycling speed than did group cyclists. The mean cycling speed was highest for those cyclists who used the road rather than the sidewalk. The mean cycling speed decreased for cyclists without helmets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".