Benchmarking Elevation Plus Land Surface Parameters Finds FathomDEM and Copernicus DEM Win as Best Global DEMs
Bibliographic record
Abstract
We evaluated six global digital elevation DEMs at 1-arc-sec resolution: CopDEM and AW3D30, which are digital surface models (DSMs), and EDTM, GEDTM, FABDEM, and FathomDEM, which are digital terrain models (DTMs). We compared them to reference DTMs created by mean aggregation from 1–2 m lidar-derived DTMs from national mapping agencies, using 1510 approximately 10 × 10 km test tiles from the United States and western Europe. Our criteria used the grids for elevation and derived land surface parameters (LSPs), including characteristics of the difference distributions and the fraction unexplained variance derived from grid correlations. The best DEM depends on the LSP used and the characteristics of the test tile, especially average slope, barrenness, and forest coverage. FathomDEM emerged as the best among the DEMs, with CopDEM the best overall for the DEMs with unrestricted licenses. GEDTM performed poorly. This is especially important for LSPs like curvature measures, which require higher-order partial derivatives for computation, and which should be used very cautiously.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".