Investigating the influence of segmentation in estimating safety performance functions for roadway sections
Bibliographic record
Abstract
Safety performance functions (SPFs) are crucial to science-based road safety management. Success in developing and applying SPFs, apart data quality and availability, depends fundamentally on two key factors: the validity of the statistical inferences for the available data and on how well the data can be organized into distinct homogeneous entities. The latter aspect plays a key role in the identification and treatment of road sections or corridors with problems related to safety. Indeed, the segmentation of a road network could be especially critical in the development of SPFs that could be used in safety management for roadway types, such as motorways (freeways in North America), which have a large number of variables that could result in very short segments if these are desired to be homogeneous. This consequence, from an analytical point of view, can be a problem when the location of crashes is not precise and when there is an overabundance of segments with zero crashes. Lengthening the segments for developing and applying SPFs can mitigate this problem, but at a sacrifice of homogeneity. This paper seeks to address this dilemma by investigating four approaches for segmentation for motorways, using sample data from Italy. The best results were obtained for the segmentation based on two curves and two tangents within a segment and with fixed length segments. The segmentation characterized by a constant value of all original variables inside each segment was the poorest approach by all measures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".