Data‐Driven Approach for Passenger Assignment in Urban Rail Transit Networks: Insights From Passenger Route Choices and Itinerary Choices
Bibliographic record
Abstract
Congestion in urban rail transit (URT) systems often results in passengers being left behind on platforms due to trains’ reaching capacity. Distinguishing between the travel choice behaviors of passengers who board the first arriving train (Type I passengers) and those who are left behind (Type II passengers) in passenger assignment is essential for effective URT passenger management. This paper proposes a data‐driven passenger‐to‐train assignment model (DPTAM) that leverages automated fare collection (AFC) data and automated vehicle location (AVL) data to differentiate between the travel choice behaviors of the two types of passengers. The model comprises two modules based on passenger travel choice behavior: the passenger route choice model (PRCM) and the passenger itinerary choice model (PICM). The PRCM employs a granular ball–based density peaks clustering (GB‐DP) algorithm to estimate passengers’ route choices based on historical data, enhancing precision and efficiency in passenger classification and route matching. The PICM incorporates tailored itinerary selection strategies that consider train capacity constraints and schedules, enabling accurate inference of passenger itineraries and localization of their spatiotemporal states. The model also estimates train loads and left‐behind probabilities to identify congested periods and sections. The effectiveness of DPTAM is validated through synthetic data, demonstrating superior assignment accuracy compared to benchmarks. Additionally, real‐world data from Chengdu Metro reveal the impact of congestion on travel behavior and effectively identify congested periods and high‐demand stations and sections, highlighting its potential to enhance URT system efficiency and passenger management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".