Development and Validation of a Crash Culpability Scoring Tool
Bibliographic record
Abstract
OBJECTIVE: Several traffic safety research techniques require researchers to separate crash-involved drivers into culpable and nonculpable. Nonculpable drivers are assumed to be randomly involved in crashes by external factors and to approximate a noncollision control population. If this is true, factors that increase crash risk should be found more often in culpable than in nonculpable drivers. Though a culpability scoring tool has been developed for research purposes, that tool does not adequately address winter driving conditions (Robertson and Drummer 1994). Moreover, traditional culpability scoring requires assessors to read and score individual collision reports. The purpose of this study is to develop and validate an automated, rule-based Canadian culpability scoring tool that is capable of rapidly scoring police crash reports from large administrative datasets. METHODS: We used an iterative approach to develop and validate our tool. First, the Robertson-Drummer culpability scoring tool was modified to include the extensive police report data collected in the British Columbia Traffic Accident System (TAS) and to account for winter driving conditions. This was done in consultation with traffic safety experts. The scoring tool was automated, employing a rule-based decision model that avoids interpretation of free-text reports. The scoring tool was applied to 73 collisions (134 drivers). Two experts also reviewed these collisions and determined the culpability of each driver. Discrepant cases were discussed to understand why the scoring tool differed from the expert assessment and the scoring tool was modified accordingly. The final tool was compared with expert assessment on another sample of 96 crashes. The tool was also applied to a sample of 2086 crash-involved drivers with known blood alcohol concentrations (BACs) and the adjusted odds of culpability were calculated for several BAC ranges. RESULTS: The final scoring tool included 7 factors and had content validity for traffic safety experts. It had excellent agreement with expert scoring on the first set of collisions (kappa = 0.83, 95% confidence interval [CI]: 0.75-0.91) and on the second set (kappa = 0.84, 95% CI: 0.77-0.92). When applied to crash-involved drivers with known BAC levels, the scoring tool exhibited predictive validity: the odds of culpability increased with higher BACs, consistent with the known dose effect of BAC on crash risk. CONCLUSIONS: We have developed an automated culpability scoring tool contextualized to Canadian driving conditions. This tool will allow road safety researchers to assess collision responsibility in large administrative data sets derived from police reports.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".