Data-Driven Bike-Share Ridership Prediction and Network Optimization
Bibliographic record
Abstract
Shared micro-mobility systems, particularly station-based bike-sharing networks, have become key components of urban transportation, yet their planning remain challenged by spatial and technological complexities. This dissertation develops integrated models for ridership prediction, station placement optimization, and electrification planning to address these challenges. First, a customized Graph Neural Network framework using GraphSAGE is introduced for station-to-station ridership prediction, integrating network topology, sociodemographic features, and station attributes. Applied to Toronto, the model outperforms linear, spatial, and tree-based benchmarks, demonstrating its ability to capture latent dependencies and support demand-responsive planning. Second, a continuum approximation model is proposed for station placement optimization, using a force-based algorithm that balances attraction from demand centers with inter-station forces. This ridership-driven approach departs from conventional accessibility methods by directly aligning locations with demand. Applied to Vancouver, the model reveals optimal spacing patterns and highlights strategies for ridership-driven network expansion under varying demand conditions. Third, the dissertation extends infrastructure planning to electrified systems by introducing a two-dimensional Markovian state-of-charge framework for e-bikes. A heuristic charger deployment algorithm, enhanced by a single-pooling state approximation, maximizes expected ridership and identifies high-impact charging locations, achieving near-optimal performance in case studies from Pittsburgh, Vancouver, and San Francisco. Finally, the models are integrated into a web-based, GIS-enabled decision-support tool that combines predictive, prescriptive, and descriptive analytics to enable scenario-based planning. Demonstrations in Toronto and Vancouver illustrate the tool’s scalability and practical value. This research advances methodological foundations and practical tools for developing resilient, data-driven, and electrified bike-sharing networks across diverse urban environments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".