Predicting urban flooding susceptibility of public transit systems using machine learning approaches
Bibliographic record
Abstract
Urban floods often cause the functional disruption of public transit systems, thereby impeding people's mobility and resulting in adverse socio-economic consequences. Climate change, rapid urbanization, and unplanned disaster management further increase trends of urban floods with higher frequency and intensity. This study employs data-driven machine learning (ML) models for predicting the flooding susceptibility of public transit systems in Toronto, ON, Canada. Four ML approaches are employed to evaluate the future risks of public transit systems being inundated by flooding events: 1) Random Forest (RF); 2) eXtreme Gradient Boosting (XGBoost); 3) K Nearest Neighbor (KNN); and 4) Naïve Bayes (NB). We estimate flooding probability based on the relationship between flood inundation events and their contributing factors. Flood-plain maps by Toronto and Region Conservation Authority (TRCA) are used to generate flood and non-flood locations as a basis for training (70% of samples) and validating (30% of samples) ML models. We use the Area-Under Receiver Operating Characteristic (AUROC) curve to evaluate the prediction results of ML models. All four models show a high level of accuracy higher than 95%. Our results demonstrate public transit systems near major river channels are highly susceptible to floods which corroborated with historical flooding incidents in 2013 and 2018. The outcome of the study can be helpful to enhance the resilience of public transit systems in the city of Toronto and can facilitate evidence-based planning and policy to make cities more sustainable, livable, and resilient against flood hazards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".