Understanding bus delay patterns under different temporal and weather conditions: A Bayesian Gaussian mixture model
Bibliographic record
Abstract
In public transit systems, bus delays significantly impact service reliability and passenger satisfaction. Causal delays, consisting of link running and stop dwell delays, are critical factors contributing to overall bus delay patterns. This paper develops a Bayesian probabilistic model to analyze bus delay patterns with a focus on causal delays under varying weather and temporal conditions, which can help to understand how the underlying causal delay patterns contribute to arrival delay patterns. Employing a Gaussian mixture model integrated with a topic model approach, the study analyzes causal delays as multivariate random variables , capturing the influence of temporal and weather conditions on bus service reliability. For model inference, we propose a Markov Chain Monte Carlo (MCMC) sampling method to estimate the model parameters. The analysis is conducted using real-world data from a bus route in Calgary, Canada. We categorize the identified delay patterns into four on-time categories: extreme earliness, moderate earliness, extreme lateness, and moderate lateness. Results indicate that adverse weather significantly influences extreme delay patterns in particular, suggesting the necessity for transit agencies to consider these factors in schedule optimization. Beyond pattern identification , the proposed model offers probabilistic delay estimation, enabling accurate forecasting of future delays based on current conditions and observations. Validation results demonstrate that our probabilistic estimates align closely with observed data, proving the model’s practical applicability in real-time operations and offering actionable insights to enhance the punctuality and efficiency of urban bus services.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".