Enriching Roadside Safety Assessments Using LiDAR Technology: Disaggregate Collision-Level Data Fusion and Analysis
Bibliographic record
Abstract
Fatalities and serious injuries still represent a significant portion of run-off-the-road (ROR) collisions on highways in North America. In order to address this issue and design safer and more forgiving roadside areas, more empirical evidence is required to understand the association between roadside elements and safety. The inability to gather that evidence has been attributed in many cases to limitations in data collection and data fusion capabilities. To help overcome such issues, this paper proposes using LiDAR datasets to extract the information required to analyze factors contributing to the severity of ROR collisions on a localized collision level. Specifically, the paper proposes a new method for extracting pole-like objects and tree canopies. Information about other roadside assets, including signposts, alignment attributes, and side slopes is also extracted from the LiDAR scans in a fully automated manner. The extracted information is then attached to individual collisions to perform a localized assessment. Logistic regression is then used to explore links between the extracted features and the severity of fixed-object collisions. The analysis is conducted on 80 km of roads from 10 different highways in Alberta, Canada. The results show that roadside attributes vary significantly for the different collisions along the 80 km analyzed, indicating the importance of utilizing LiDAR to extract such features on a disaggregate collision level. The regression results show that the steepness of side slopes and the offset of roadside objects had the most significant impacts on the severity of fixed-object collisions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".