Factors Affecting Classification of Road Segments into High- and Low-Speed Collision Regimes
Bibliographic record
Abstract
The safety of locations operating under high-speed conditions could significantly differ from that of locations operating under low-speed conditions. Therefore, different approaches must be adopted when speed and safety are analyzed and managed at locations operating under different regimes. However, it is necessary first to understand the factors affecting the speed–collision classification of a site. Locations operating under high speeds are typically expected to have more collisions compared with locations in which speeds are low. Some locations, however, might experience a high collision rate even when speeds are low, or vice versa. This study aimed to identify the factors that affected the site classification into any of those categories by using data collected on roads in Edmonton, Alberta, Canada. Locations were divided into four speed–collision bins (high collision, high speed; high collision, low speed; low collision, high speed; low collision, low speed), and geographic information system maps of locations were produced to explore the spatial distribution of those locations. Moreover, logistic regression was used to understand the role of different factors in identifying the speed–collision bin to which a certain location belonged. The results reveal that locations with high collision rates but low speeds have a relatively high population of heavy vehicles and trucks as well as high speed variability. As for locations with low collision rates and high speeds, these sites were found to have a high level of protection through the presence of medians and shoulders with relatively low access density.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".