Classification of Moral Decision Making in Autonomous Driving: Efficacy of Boosting Procedures
Bibliographic record
Abstract
Autonomous vehicles (AVs) face critical decisions in pedestrian interactions, necessitating ethical considerations such as minimizing harm and prioritizing human life. This study investigates machine learning models to predict human decision making in simulated driving scenarios under varying pedestrian configurations and time constraints. Data were collected from 204 participants across 12 unique simulated driving scenarios, categorized into young (24.7 ± 3.5 years, 38 males, 64 females) and older (71.0 ± 5.7 years, 59 males, 43 females) age groups. Participants’ binary decisions to maintain or change lanes were recorded. Traditional logistic regression models exhibited high precision but consistently low recall, struggling to identify true positive instances requiring intervention. In contrast, the AdaBoost algorithm demonstrated superior accuracy and discriminatory power. Confusion matrix analysis revealed AdaBoost’s ability to achieve high true positive rates (up to 96%) while effectively managing false positives and negatives, even under 1 s time constraints. Learning curve analysis confirmed robust learning without overfitting. AdaBoost consistently outperformed logistic regression, with AUC-ROC values ranging from 0.82 to 0.96. It exhibited strong generalization, with validation accuracy approaching 0.8, underscoring its potential for reliable real-world AV deployment. By consistently identifying critical instances while minimizing errors, AdaBoost can prioritize human safety and align with ethical frameworks essential for responsible AV adoption.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".