Data-driven modeling method for analyzing grade crossing safety
Bibliographic record
Abstract
A grade crossing is defined as an intersection between a roadway and a railway at the same elevation or grade. Multiple new prevention measures have been implemented to reduce the number of train-vehicle collisions; however, crossing safety is still a major issue as accidents still frequently occur. The push for data-driven models to evaluate risks at grade crossings has also increased to keep up with the changing technologies. There are many different protection types (gates with bells, cross-buck, stop-sign, mirrors and etc.) that serve to warn or stop oncoming traffic. Many attributes have an inherent impact on accident frequency; including the protection type, train speed, traffic volume and e.t.c. To address which factors are most important, we propose a data-driven modeling method to effectively analyze the impact of multiple factors that affect crossing safety and subsequently provide scientific insight for key factors for enhancing crossing safety. In this work, the Canadian crossing accident database for the years of 2004 – 2013 was used with additional generated features to enhance the scope of the study. These include features that were computed using GIS and sightline measurements. Data-driven modeling using RandomForests were used to rank and analyze 21 attributes for each protection type. From the analysis results it is possible to identify which key factors have the highest influence on improving safety and collision prediction at grade crossings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".