Physically Based Dimensionless Features for Pluvial Flood Mapping With Machine Learning
Bibliographic record
Abstract
Abstract Rapid delineation of flash flood extents is critical to mobilize emergency resources and to manage evacuations, thereby saving lives and property. Machine learning (ML) provides a promising solution for this rapid delineation, offering a computationally efficient alternative to high‐resolution 2D flood models. However, even when trained on diverse geographic regions, ML models typically require retraining to perform well in new locations, and therefore often fail to generalize to never‐before‐seen conditions. To improve ML generalization, we apply Buckingham theorem to derive dimensionless terms across multiple spatial scales. These multiscale terms represent ratios of the relevant physical quantities governing the flooding process. Since the scaling laws of these dimensionless terms encode process similarity across physical scales, these terms enhance ML transferability to unseen locations. This is demonstrated by incorporating them as features in a logistic regression model for delineating flood extents. The features were calculated at different scales by varying accumulation thresholds for stream delineation. The ML flood maps, with an average AUC of 0.89, compared well with the results of 2D hydraulic models that are the basis of the Federal Emergency Management Agency flood hazard maps. The dimensionless features outperformed dimensional features, with some of the largest gains in the AUC (of 20%) occurring when the model was trained in one region and tested in another. Dimensionless and multi‐scale features in ML flood modeling have the potential to improve generalization, enabling mapping in unmapped areas and across a broader spectrum of landscapes, climates, and events.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".