Scene-Centric Vehicle Trajectory Prediction at Cooperative Intersection Using Decision-Aware Attention Graph Transformer
Bibliographic record
Abstract
Roadside sensors offer a fixed, unobstructed vantage point that can overcome line-of-sight limitations in autonomous driving environments by sharing critical perception data with nearby road agents. While this cooperative approach enhances situational awareness, it also introduces significant computational and communication overhead for autonomous vehicles (AVs). To address this challenge, we propose the Heterogeneous Decision-Aware Attention Graph Transformer (HDAAGT)—a non-autoregressive, encoder-only transformer architecture designed for real-time vehicle trajectory prediction. HDAAGT processes detection data from roadside infrastructure to forecast future vehicle trajectories and communicates these predictions to surrounding agents. By offloading intensive computations from AVs and minimizing transmission latency, our approach improves responsiveness and enables more efficient cooperative perception at intersections and other complex driving scenarios. HDAAGT integrates lane positioning, traffic light states, and vehicle kinematics, enabling a decision-aware graph attention mechanism that models agent-agent and agent-environment interactions. By leveraging a fisheye-based detection and tracking pipeline, our approach eliminates the need for multiple cameras and enables HDAAGT to generate reliable trajectory predictions across the full intersection. We validate our model on the Fisheye-MARC and SinD datasets, demonstrating the capability of HDAAGT in predicting vehicle motion in complex urban intersections with a 1.28 m final displacement error. Additionally, we introduce a new 31k-frame fisheye intersection dataset, the largest of its kind in object tracking, to advance research in intersection-based trajectory prediction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".