Discovering Insightful Rules among Truck Crash Characteristics using Apriori Algorithm
Bibliographic record
Abstract
This study aims to discover hidden patterns and potential relationships in risk factors in freight truck crash data. Existing studies mainly used parametric models to analyze the causes of freight vehicle crashes. However, predetermined assumptions and underlying relationships between independent and dependent variables have been cited as its limitations. To overcome these limitations and provide a better understanding of factors that lead to truck crashes on the expressways, we applied the Association Rules Mining (ARM) technique, which is a nonparametric method. ARM quantifies the interrelationships between the antecedents and consequents of truck-involved crashes and provides researchers with the most influential set of factors that leads to crashes. We utilized a freight vehicle-involved crash data consisting of 19,038 crashes that occurred on the Korean expressways from 2008 to 2017 for this investigation. From the data, 90,951 association rules were generated through ARM employing the Apriori algorithm. The lift values estimated by the Apriori algorithm showed the strength of association between risk factors, and based on the estimated lift values, we identified key crash contributory factors that lead to truck-involved crashes at various segment types, under different weather conditions, considering the driver’s age, crash type, driver’s faults, vehicle size, and roadway geometry type. From the generated rules, we demonstrated that overspeeding with medium-weight trucks was highly associated with crashes during the rainy weather, whereas drowsy driving during the evening was correlated with crashes during fine weather. Segment-related crashes were mainly associated with driver’s faults and roadway geometry. Our results present useful insights and suggestions that can be used by transport stakeholders, including policymakers and researchers, to create relevant policies that will help reduce freight truck crashes on the expressways.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".