Determining the validity of the AMA guide: A retrospective analysis of the Assessment of Driving Related Skills and crash rate among older drivers
Bibliographic record
Abstract
Background: Chronic health conditions associated with aging can lead to changes in driving ability. The Canadian Driving Research Initiative for Vehicular Safety in the Elderly (Candrive II) is a 5-year prospective study funded by the Canadian Institutes of Health Research aiming to develop an in-office screening tool that will help clinicians identify potentially at-risk older drivers. Currently, no tools exist to directly predict the risk of motor vehicle collision (MVC) in this population. The American Medical Association (AMA), in collaboration with the National Highway Traffic Safety Association, has designed an opinion-based guide for assessing medical fitness to drive in older adults and recommends that physicians use the Assessment of Driving Related Skills (ADReS) as a test battery to measure vision, cognition and motor/somatosensory functions related to driving. The ADReS consists of the Snellen visual acuity test, visual fields by confrontation test, Trail Making Test part B, clock drawing test, Rapid Pace Walk test, and manual tests of range of motion and motor strength. We used baseline data from the Candrive II/Ozcandrive common cohort of older drivers to evaluate the validity of the ADReS subtests. We hypothesized that participants who crashed in the 2 years before the baseline assessment would have poorer scores on the ADReS subtests than participants who had not crashed. Methods: In the Candrive II/Ozcandrive study, 1230 participants aged 70 years or older were recruited from 7 Canadian cities, 1 Australian city and 1 New Zealand city, all of whom completed a comprehensive clinical assessment at study entry. The assessment included all tests selected as part of the ADReS. Data on crashes that occurred within 2 years preceding the baseline assessment were obtained from the respective licensing jurisdictions. Those who crashed were compared to those who had not crashed on their ADReS subtest scores using Pearson’s chi-squared test and Student’s t-test. Results: Sixty-three of the 1230 participants (5.1%) were involved in an MVC within the 2 years preceding the baseline assessment. Contrary to what was expected based on the AMA guide, there were no statistically significant associations between abnormal performance on the tests constituting the ADReS and history of crash (p > 0.01). Discussion: Although limitations are inherent in a retrospective analysis, we found that abnormalities on the subtests comprising the ADReS were not associated with a recent history of crash. This suggests the need for more sensitive tools to properly assess crash risk in older drivers, for prospective analyses of risk over time and for an evidence base to support influential clinical practice guidelines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".