Evaluating the Spatial Generalizability of ML- and DL-Based Surrogate Models for Flood Depth Prediction
Bibliographic record
Abstract
Two-dimensional hydrodynamic models are widely used for flood modeling; however, their computational complexity limits their application for real-time flood forecasting and iterative frameworks requiring a large number of model runs. To address this, previous research has focused on developing surrogate models using machine learning (ML) and deep learning (DL) techniques to predict flood depth. Despite advancements, many of these models lack spatial generalizability and are constrained to the specific locations where they were trained. This study compares the performance of four surrogate models developed using three traditional ML methods (Random Forest, XG-Boost, and Least-Squares Support Vector Machine), which do not inherently account for spatial relationships and a DL method (U-Net) to evaluate their generalizability to unseen locations for identical rainfall hyetograph. The dataset used for this study was generated using a calibrated HEC-RAS flood model for Montreal Island. To enhance model performance and capture relationship between spatial characteristics and flood depth, the modeling framework incorporates multiple explanatory variables: depth to water sinks, curvature, flow accumulation, slope, elevation difference between pixel and focal mean, roughness index, topographic position index, topographic wetness index, and surface elevation. Results demonstrate superior performance of the DL-based method compared to the traditional ML approaches considered, attributed to its capacity to capture the spatial correlation of flood depths between neighboring cells. The performance of the models over unseen locations show root mean squared error (RMSE, in m) and mean absolute error (MAE, in m) of 0.336 and 0.184 for RF, 0.341 and 0.181 for XG-Boost, 0.336 and 0.183 for LS-SVM, and 0.197 and 0.105 for U-Net models, respectively. These findings are consistent with previous studies that highlight the challenges of achieving spatial generalizability in surrogate models and show the competitive accuracy of the U-Net model. While the DL-based surrogate model exhibits limitations in accurately predicting high flood depths, which are critical for flood-induced damage assessment, these results underscore both the potential of DL-based surrogate models for efficient and spatially transferable flood modeling and the need for further research to improve predictions of extreme flood depths and extend the model’s generalizability to unseen hyetographs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".