Predicting Driver Injury Severity in Single-Vehicle and Two-Vehicle Crashes with Boosted Regression Trees
Notice bibliographique
Résumé
The boosted regression tree model is an emerging nonparametric tree-based model that can capture nonlinear effects of both discrete and continuous variables without preprocessing data. The model is particularly advantageous to predict severe injuries, which are more difficult to classify because of their small amount compared with nonsevere injuries. The objectives of this study were to investigate driver injury severity with the boosted regression tree model and other nonparametric models—the classification and regression tree and Random Forests—and to evaluate performance of the boosted regression tree model in comparison with the classification and regression tree model. The study identified important factors affecting injury severity by using 5-year crash records for provincial highways in Ontario, Canada. The results of the boosted regression tree model showed that ejection from a vehicle and head-on collisions commonly had a strong association with driver injury severity. Results also showed that marginal effects of continuous variables including truck percentage, annual average daily traffic (AADT), driver age, and vehicle age on injury severity were nonlinear. In particular, their effects on the injuries of heavy-truck drivers had different patterns compared with the effects on passenger-car and light-truck drivers; the risk of severe injury to heavy-truck drivers increased as the truck percentage and AADT increased and the driver's age decreased. The boosted regression tree model predicted driver injury severity more accurately than the classification and regression tree model for both single-vehicle and two-vehicle crashes. Thus, it is recommended that the boosted regression tree model be applied with separate data sets for single-vehicle crashes and different types of two-vehicle crashes for more accurate prediction of crash injury severity.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,011 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».