MétaCan
Menu
← Back to cohort
Record W4408428954 · doi:10.5194/egusphere-egu25-14649

Evaluating the Spatial Generalizability of ML- and DL-Based Surrogate Models for Flood Depth Prediction

2025· preprint· en· W4408428954 on OpenAlexaffabout
Oveys Ziya, Laxmi Sushama, Husham Almansour

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldEnvironmental Science
TopicHydrology and Watershed Management Studies
Canadian institutionsNational Research Council CanadaMcGill University
Fundersnot available
KeywordsGeneralizability theoryFlood mythSurrogate modelSurrogate endpointComputer scienceEnvironmental scienceStatisticsGeographyMathematicsMachine learningArchaeologyMedicine

Abstract

fetched live from OpenAlex

Two-dimensional hydrodynamic models are widely used for flood modeling; however, their computational complexity limits their application for real-time flood forecasting and iterative frameworks requiring a large number of model runs. To address this, previous research has focused on developing surrogate models using machine learning (ML) and deep learning (DL) techniques to predict flood depth. Despite advancements, many of these models lack spatial generalizability and are constrained to the specific locations where they were trained. This study compares the performance of four surrogate models developed using three traditional ML methods (Random Forest, XG-Boost, and Least-Squares Support Vector Machine), which do not inherently account for spatial relationships and a DL method (U-Net) to evaluate their generalizability to unseen locations for identical rainfall hyetograph. The dataset used for this study was generated using a calibrated HEC-RAS flood model for Montreal Island. To enhance model performance and capture relationship between spatial characteristics and flood depth, the modeling framework incorporates multiple explanatory variables: depth to water sinks, curvature, flow accumulation, slope, elevation difference between pixel and focal mean, roughness index, topographic position index, topographic wetness index, and surface elevation. Results demonstrate superior performance of the DL-based method compared to the traditional ML approaches considered, attributed to its capacity to capture the spatial correlation of flood depths between neighboring cells. The performance of the models over unseen locations show root mean squared error (RMSE, in m) and mean absolute error (MAE, in m) of 0.336 and 0.184 for RF, 0.341 and 0.181 for XG-Boost, 0.336 and 0.183 for LS-SVM, and 0.197 and 0.105 for U-Net models, respectively. These findings are consistent with previous studies that highlight the challenges of achieving spatial generalizability in surrogate models and show the competitive accuracy of the U-Net model. While the DL-based surrogate model exhibits limitations in accurately predicting high flood depths, which are critical for flood-induced damage assessment, these results underscore both the potential of DL-based surrogate models for efficient and spatially transferable flood modeling and the need for further research to improve predictions of extreme flood depths and extend the model’s generalizability to unseen hyetographs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.012
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.022
Threshold uncertainty score0.043

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.012
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.066
GPT teacher head0.319
Teacher spread0.254 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes2
Has abstractyes

Explore more

Same topicHydrology and Watershed Management Studies→French-language works237,207→