Predicting Pedestrian Crossing Intentions in Adverse Weather With Self-Attention Models
Bibliographic record
Abstract
The enhancement of the vehicle perception model represents a crucial undertaking in the successful integration of assisted and automated vehicle driving. By enhancing the perceptual capabilities of the model to accurately anticipate the actions of vulnerable road users, the overall driving experience can be significantly improved, ensuring higher levels of safety. Existing research efforts focusing on the prediction of pedestrians’ crossing intentions have predominantly relied on vision-based deep learning models. However, these models continue to exhibit shortcomings in terms of robustness when faced with adverse weather conditions and domain adaptation challenges. Furthermore, little attention has been given to evaluating the real-time performance of these models. To address these aforementioned limitations, this study introduces an innovative framework for pedestrian crossing intention prediction. The framework incorporates an image enhancement pipeline, which enables the detection and rectification of various defects that may arise during unfavorable weather conditions. Subsequently, a transformer-based network, featuring a self-attention mechanism, is employed to predict the crossing intentions of target pedestrians. This augmentation enhances the model’s resilience and accuracy in classification tasks. Through evaluation on the Joint Attention in Autonomous Driving (JAAD) dataset, our framework attains state-of-the-art performance while maintaining a notably low inference time. Moreover, a deployment environment is established to assess the real-time performance of the model. The results of this evaluation demonstrate that our approach exhibits the shortest model inference time and the lowest end-to-end prediction time, accounting for the processing duration of the selected inputs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".