Comparison of climate datasets for lumped hydrological modeling over the continental United States
Bibliographic record
Abstract
Climate data measured by weather stations are crucially important and regularly used in hydrologic modeling. However, they are not always available due to the low spatial density and short record history of many station networks. To overcome these limitations, gridded datasets have become increasingly available. They have excellent continuous spatial coverage and no missing data. However, these datasets are usually interpolated using station data, with little new information besides elevation. Furthermore, minimal validation has been done on most of these datasets. This study compares three such datasets covering the continental United States to evaluate their differences and their impact on lumped hydrological modeling. Three daily time step gridded datasets with resolutions varying between 0.25° and 1 km were used in this study – Santa-Clara, Daymet and CPC. The hydrological modeling evaluation of these datasets was performed over 424 basins from the MOPEX database. Results show that there are significant differences between the datasets, even though they were essentially all interpolated from almost the same climate databases. Despite those differences, the hydrological model used in this study was able to perform equally well after a specific calibration to each dataset. While there were a few exceptions, by and large, Nash–Sutcliffe efficiency metrics obtained in validation were not statistically different from one database to the other for most basins. It appears that there are no reasons to favor one dataset versus another for lumped hydrological modeling, and that these datasets perform just as well as using the original station data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".