Exploring the factors affecting injury severity in highway and non-highway crashes in Bangladesh applying machine learning and SHAP
Bibliographic record
Abstract
To create effective preventive measures and targeted interventions, it is crucial to comprehend the contributing factors to the crash and quantify how they affect the injury, especially in least-developed countries. However, highway and non-highway crashes are linked to having distinguished characteristics, road-specific interventions, and data granularity. Combining all sorts of crashes into a single model may offer fewer insights than one would anticipate when building safety countermeasures. This research compares CART, RF, GBM, XGBoost, LightGBM, CatBoost, and AdaBoost and effectively simulates the complex relationship between collision injury severity and risk factors for both highway and non-highway crashes. Additionally, the Shapley Additive exPlanation (SHAP) framework is presented to explain the contribution of each risk factor from the output of the most appropriate classifier, thereby assisting in the construction of safety countermeasures and crash modification factors. GBM classifier was found to be the best classifier in terms of G-mean and AUC scores for both highway and non-highway models. Global SHAP values show that the type of collision, followed by the vehicle type, the vehicle involved, and road division, are the highest contributing factors for injury severity in highway crashes. For injury severity in non-highway crashes, the most important factors are the type of collision, followed by road division, vehicle type, and location type. Policy implications based on the study's findings have been suggested to develop successful preventive strategies and focused interventions. The study concludes by discussing the scope of future studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".