Assessing the Field Relevance of Testing protocols and Injury risk functions employed in new car assessment programs
Bibliographic record
Abstract
Over the past two decades the popularity of consumer crash test programs, commonly referred to as New Car Assessment Programs (NCAP), has grown across the world. They are popular among government regulators as they afford a means of promoting safety innovations and levels of vehicle performance beyond those dictated by national standards. They also fulfill the demand for information regarding the safety ranking of vehicles among consumers contemplating the purchase of a new vehicle. There is no question that consumer crash test programs greatly influence vehicle design changes as well as accelerate the fitment of new safety features. The extent to which these changes can be expected to reduce serious and potentially fatal injuries will be influenced by how well the testing protocols and associated rating schemes correctly reflect the nature of the residual safety problem they seek to address. Drawing on data contained primarily in the US National Automotive Sampling System (NASS), the field relevance of current and proposed testing and rating protocols addressing frontal crash test protection is examined. Emphasis is placed on examining how accurately injury rates computed from the dummy responses measured in consumer crash tests correspond to actual injury rates observed in the field. Additional data from Canadian field investigations and US databases such as the National Motor Vehicle Crash Causation Survey (NMVCCS) are examined to see how well frontal airbag firing times, crush pulse durations and other determinants of injury are replicated in consumer testing protocols. This portion of the analysis draws on data obtained from Event Data Recorders (EDR) in both field collisions and staged tests of the same vehicle model. Vehicle rankings and overall frontal crash test ratings were found to be particularly sensitive to the choice of injury risk functions employed in the test. This was particularly true in the case of injury risk functions used to assess neck injury potential. Neck injury risk derived from Nij was found to show the least agreement with the field. Agreement between field chest injury rates and those derived from crash tests was improved considerably when chest injury risk functions for older occupants were employed. The paper concludes with a discussion of how different current testing protocols could be improved to enhance their field relevance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.386 | 0.640 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".