Road Collision Analysis and Prediction Using Machine Learning Approaches
Bibliographic record
Abstract
Road travel accounts for most traffic accidents worldwide. Improvements in road safety, education, recent technology advancements, and other environmental factors have decreased the number of collisions in developed nations. Many countries, provincial, and local governments envision the possibility of zero fatalities or serious injuries in the near future. Thus, it is essential to develop road traffic accident prediction models to support such a vision. On the one hand, classical statistical models have been applied to develop prediction models throughout the literature. These models provide interpretable parameters at the expense of poor generalization when faced with complex and nonlinear relationships. On the other hand, data-driven methods utilizing Machine Learning (ML) approaches have been used recently to deal with the drawbacks of classical models, which showed promising results. Road accidents result from many factors, including spatial, temporal and external factors. Those factors may influence the occurrence of accidents differently, according to the location and time of accidents. Thus, it is essential to consider the area-specific influential factors while analyzing and developing prediction models. Canada is the second coldest country globally, and its extreme weather has a higher effect on accidents than the other countries, which must be addressed. This thesis seeks to explore determinants of road collisions, emphasizing Canadian weather. It then compares classical and ML models for collision prediction. Furthermore, it introduces the most influential factors in crashes with respect to Calgary's weather. All study parts are performed on the collisions data in Calgary, Alberta, Canada, between 2017 to 2020. It is shown that all the weather attributes are correlated to collisions. It shows the importance of considering the weather attributes in accident analysis and prediction. Based on the nature of the collisions dataset, which is tabular and heterogeneous, Neural Networks showed higher performances than the other investigated model, with 92% accuracy. The proposed models can be used for policy-making and individual usage in Canadian cities since the effect of all the weather features is already embedded in the models. In order to demonstrate the thesis's applicability, a new speed limit is recommended utilizing the developed models for Deerfoot TR SE. Results showed, for instance, if the speed limit is decreased from 100 to 90 km/h on Deerfoot TR SE, a 5% accident reduction is predicted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".