Data Analytics for Dependable Transportation Systems in a Smart City
Bibliographic record
Abstract
Public transit is an important component of the day-to-day activities of many people. It provides a cost-effective and convenient way for individuals to commute to work, school, and other destinations. Bus transit is a vital mode of transportation for students, as it enables them to commute to and from their educational institutions. Delays in bus schedules can have severe consequences-such as missing exams, meetings, and other important engagements-in daily activities of city residents. Hence, in this paper, we present a data science solution for mining and transportation analytics on public transit on-time performance data. Knowledge discovered from these data helps improve public transit performance, and thus enhance rider experience in a city. This helps build a smart city. To elaborate, our solution adapts frequent pattern mining, which identify uncover variations in transit performance across various neighborhoods. Through identifying significant findings, we establish correlations to determine the factors contributing to bus delays in specific areas. Improving the bus arrival or departure time can have a positive impact on the overall usability and attractiveness of bus transit for commuters since people are more likely to use buses when they can rely on them to arrive on time and get to their destinations promptly. Our solution also provides users features to visualize the discovered knowledge about the bus departure time in all the neighborhoods at different times of the day. Evaluation results on real-life public transit data from a Canadian city demonstrated the practicality of our data science solution towards the building of a dependable transportation system in a smart city.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".