Integrating Information From Prior Research Into a Before-after Road Safety Evaluation Through Bayesian Approach and Data Sampling
Bibliographic record
Abstract
Before-after road safety evaluation (B/A) to measure safety treatment effect is a key mission in road safety management, and has fueled considerable research. However, previous research in this area has been overwhelmingly dedicated to safety model estimation with less emphasis on other methodological issues. As a result, there continues to be uncertainty in the validity of treatment effect estimates. This study seeks, with innovative paradigms, a systematic solution by solidifying methodologies for every essential step of a thorough B/A process to secure its ultimate validity. Methodologies of data sampling and processing, and before and after model development, both vital procedures that have been historically neglected, are investigated. A pre-test data sampling approach to select reference groups is established in the context of B/A application. A post-assignment propensity score matching method is developed in order to further eliminate statistical bias while the treatment effect indicator – collision reduction ratio (CRR) – is being estimated. Rather than focus on single safety model development as is common in traffic safety research, this study seeks all viable knowledge by employing various safety measures including collision and safety surrogates, by embedding several adaptable random distributions, by fitting models through both "Frequentist" and "Bayesian" approaches, and by exploring a variety of model forms and components. Accordingly, the output of this study is not a "best" single model, but rather an amalgamation of diversified models. The diversity is shown to be attractive in terms of information conveyed, especially for the B/A process. Finally, this study succeeds in finding a methodology to integrate all of the diverse knowledge sources. The Bayesian Model Averaging (BMA) method is investigated and developed to integrate a variety of statistical significant models without exclusion, in forging a unified model. All methodologies explored and developed in this study are essential to secure the validity of the B/A process. As important, they are substantially connected to each other. Should one method be deficient, the remaining steps cannot guarantee validity of B/A process. As a whole, these methodologies, if properly developed and applied, constitute a logical chain to estimate treatment effect with minimal errors and high validity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.110 | 0.195 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.005 | 0.007 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".