Adaptive Transit Signal Priority Algorithms for Optimizing Bus Reliability and Travel Time using Deep Reinforcement Learning
Bibliographic record
Abstract
Transit Signal Priority (TSP), a broadly used traffic signal control strategy, is traditionally designed for reducing transit delays at signalized intersections. Conditional TSP is popular in field applications, and research studies have often designed such strategy based on criteria-oriented yes or no rules, which target simple objectives and do not guarantee optimality. Although recent objective-oriented TSP systems began to consider the optimization of more objectives, headway adherence is rarely included. TSPs that addressed transit reliability issues commonly focused on improving the schedule adherence and were only able to reduce schedule delays by expediting buses. Headway regularity which is a critical performance indicator for high-frequency services has not received much attention. Algorithms that expedite late buses only have limited capacity to resolve short headway gaps. Moreover, objective-oriented TSPs most frequently use mathematical programming methods, the main concern with which is the requirement of explicit formulation and representation of the system performance, which typically includes assumptions that simplify the dynamic traffic environment greatly. These assumptions, like deterministic traffic flow could ignore or oversimplify the stochastic characteristics of the system. This PhD thesis proposes dual-objective adaptive TSP algorithms optimized using Deep Reinforcement Learning (DRL). These TSPs use loop detectors, and they optimize transit delays and reliability (i.e., headway adherence) at an individual intersection or multiple intersections. The proposed algorithms are trained and tested in a stochastic microsimulation environment in Aimsun Next that models a transit line segment with reliability issues in the City of Toronto. The performance of developed TSPs is compared against carefully developed baseline scenarios, including background signal timing plans without TSP, current TSP algorithm used in the field in the City of Toronto, and more advanced TSP algorithms with a machine-learning based bus arrival prediction model or DRL agents. These TSP systems are evaluated on a series of aspects including bus headway adherence, percentage of extreme headways, travel time, passenger experience, effectiveness under different traffic demand levels, and impact on the cross-street traffic delays. The developed DRL-based TSPs provide noticeable improvement in headway adherence and travel time at the individual and multiple intersections levels.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".