MétaCan
Menu
Retour à la cohorte
Enregistrement W2991313386 · doi:10.1177/1071181319631481

Collision Risk Assessment Using Naturalistic Data from a Rent-A-Car Fleet

2019· article· en· W2991313386 sur OpenAlexaff
Dongsoon Min, Trevor Waite, Birsen Donmez

Notice bibliographique

RevueProceedings of the Human Factors and Ergonomics Society Annual Meeting · 2019
Typearticle
Langueen
DomaineEngineering
ThématiqueTraffic and Road Safety
Établissements canadiensUniversity of Toronto
Organismes subventionnairesnon disponible
Mots-clésCollisionBluetoothComputer scienceGlobal Positioning SystemTransport engineeringComputer securityEngineeringTelecommunicationsWireless

Résumé

récupéré en direct d'OpenAlex

We analyzed commercial fleet operations data collected by a South Korean rent-a-car company, SK Networks Co. Ltd., to evaluate the differences between collision-free and collision-involved drivers with the ultimate goal of predicting driver collision risk. The first objective was to identify critical variables related to collision risk. The second objective was to build and compare classification models to predict the colli-sion involvement of a driver. Data used in the analysis were collected through Long-Range (LoRa) Internet of Things (IoT) modem-Fleet management system (FMS) devices, a first commercial implementation of LoRa modems in the vehicle. These devices have five main built-in modules, i.e., On-Board Diagnostics (OBD-II) Connector, GPS, LoRa modem, Gravity sensor, and Bluetooth. They can communicate with the vehi-cle, the driver’s smartphone, and the host server. Data from 3,854 drivers with a total of 2.19 million trips recorded in 2018 were explored. Out of these 3,854 drivers, 514 (13.3%) were involved in at least one collision. Predictor variables were selected based on previous research that uti-lized naturalistic data to identify factors affecting collision risk (Dingus et al., 2016; Tselentis, Yannis, & Vlahogianni, 2016; Bian, Yang, Zhao, & Liang, 2018; Jin et al., 2018). Forty-eight predictor variables that may affect collision risk were selected, which can be categorized into two groups: 27 variables were related to the business and the environment, characterized by how much drivers traveled, in what type of vehicles, on what types of roads, and during what times of the day; the other 22 variables were driving behavior-related variables, capturing overspeeding, potential fatigue, rapid speed changes, and counts of traffic regulation violations. After a feature selection phase based on univariate analysis, nine variables were select-ed to be used in the classification models. These selected vari-ables are running driving time (driving time excluding idling time), trip frequency per thousand kilometer driving, accumu-lated count of violation, accumulated amount of fine, the per-centage of trips driving a compact car (<1,000 cc), the per-centage of trips driving older than a 2016 car model, the per-centage of trips during 6 a.m.-9 a.m., the percentage of trips that ended during 2 a.m.-7 a.m., and the sum of rapid accelera-tion and deceleration frequencies per kilometer. A total of twenty classification models were built and compared to classify collision-involved and non-collision in-volved drivers: 5 classification modeling techniques (Logistic Regression (LR), Random Forest (RF), k-Nearest Neighbor (kNN), Support Vector Machine (SVM) and Gradient boosted trees (GBT)) x 4 sampling methods (Up, Down, Smote, and No-sampling). The GBT-down sampled model showed the best classification performance according to Area under the Curve (0.804) and Area under the Precision and Recall Curve (0.406) statistics. Comparing relative variable importance val-ues for the best three classification models (GBT, RF, and LR), both running driving time and violation count were found to be the most influential variables, followed by the sum of rapid acceleration and deceleration frequencies, accumulated amount of fine, trip frequency per thousand kilometer driving, and the percentage of trips driving a compact car. These re-sults agree with the results of previous naturalistic studies: driver behavior-related variables are highly related to collision likelihood, although running driving time in our dataset was likely dictated by businesses. This dataset provided us with a unique opportunity to take an in-depth look at the relationship between collisions and business, environment, and driving behavior-related variables by using naturalistic data from newly-invented LoRa IoT-FMS devices. To the best of our knowledge, this study was the first naturalistic study connecting both driving data and various types of traffic violations (e.g., overspeeding, lane, sign, park-ing, toll fees, and fine amount). Interestingly, non-driving-related violation types such as parking or toll-fee violation counts were also strongly correlated with collision involve-ment; suggesting that collision-involvement is likely not just a skill issue but also an attitude issue regarding the law. In terms of industrial applications, this study suggests multiple oppor-tunities. Through a better understanding of the influential vari-ables related to collision-involvement (e.g., accumulated vio-lations), fleet operators can build policies to enhance their fleet safety, reducing collision rates and the associated costs. Further, in the long-term, this study can provide a framework for developing a Usage-Based Rent-a-car (UBR) service for car rental field, similar to Usage-Based Insurance (UBI), which can reduce drivers’ rental fees based on their driving behaviors.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,003
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,032
Score d'incertitude au seuil0,063

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0010,003
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,001
Bibliométrie0,0020,001
Études des sciences et des technologies0,0000,000
Communication savante0,0000,001
Science ouverte0,0010,001
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,021
Tête enseignante GPT0,248
Écart entre enseignants0,227 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2019
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueProceedings of the Human Factors and Ergonomics Society Annual MeetingMême sujetTraffic and Road SafetyTravaux en français237 207