Predicting Driver Injury Severity in Single-Vehicle and Two-Vehicle Crashes with Boosted Regression Trees
Bibliographic record
Abstract
The boosted regression tree model is an emerging nonparametric tree-based model that can capture nonlinear effects of both discrete and continuous variables without preprocessing data. The model is particularly advantageous to predict severe injuries, which are more difficult to classify because of their small amount compared with nonsevere injuries. The objectives of this study were to investigate driver injury severity with the boosted regression tree model and other nonparametric models—the classification and regression tree and Random Forests—and to evaluate performance of the boosted regression tree model in comparison with the classification and regression tree model. The study identified important factors affecting injury severity by using 5-year crash records for provincial highways in Ontario, Canada. The results of the boosted regression tree model showed that ejection from a vehicle and head-on collisions commonly had a strong association with driver injury severity. Results also showed that marginal effects of continuous variables including truck percentage, annual average daily traffic (AADT), driver age, and vehicle age on injury severity were nonlinear. In particular, their effects on the injuries of heavy-truck drivers had different patterns compared with the effects on passenger-car and light-truck drivers; the risk of severe injury to heavy-truck drivers increased as the truck percentage and AADT increased and the driver's age decreased. The boosted regression tree model predicted driver injury severity more accurately than the classification and regression tree model for both single-vehicle and two-vehicle crashes. Thus, it is recommended that the boosted regression tree model be applied with separate data sets for single-vehicle crashes and different types of two-vehicle crashes for more accurate prediction of crash injury severity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".