Bibliographic record
Abstract
Determining the various components of road travel that represent elevated levels of risk would be a fundamental step in the identification of risk issues on our roads. Once identified, the occurrence of these road safety issues can be targeted for improvements, thus making progress towards the safety on our roads. To that end, we have developed a risk analysis model and have determined some road safety issues of high-risk. The model weighs incident data from theNational Collision Data Base (NCDB) against exposure data from the Canadian Vehicle Survey(CVS). Fatality rates and relative risk levels were computed for all groupings of data common to both datasets. Without the exposure data (CVS), it would only be possible to compute casualty rates by population, number of licensed drivers, or the number of registered vehicles. This is limited; it does not allow one to compare risk factors at micro levels; for example, the risk of nighttime driving versus daytime driving. The exposure data enables the calculation of standardized risk values based on the actual amount of kilometers driven during the respective times, and allows for the comparison of the different risk factors to one another. There are two ways of calculating risk; casualty rates and relative risk. Casualty rates take the number of road incidents that occurred under certain conditions divided by the number of kilometers driven under the same travel conditions. One could calculate similar rates based on different conditions in order to make comparisons. The relative risk is simply the proportion of collisions with a certain characteristic versus the proportion of road travel with the same characteristic. It has no units, and if the risk is greater than 1, it means that the collisions with the characteristic under study are over-represented relative to the kilometers driven having the same characteristic. For example, a relative risk of 2 means that the proportion of incidents with the characteristic under study is twice the corresponding proportion of kilometers driven. As one can imagine, this allows us to extend our analysis to implemented programs, such as graduated licensing programs, in order to determine their effectiveness on reducing risk on our roads.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".