A harmonized 2000–2024 dataset of daily river ice concentration and annual phenology for major Arctic rivers
Bibliographic record
Abstract
Abstract. River ice plays a critical role in Arctic freshwater routing, navigation safety, and biogeochemical exchange. However, consistent, daily-resolved observations across the pan-Arctic remain scarce. Here we present a harmonized, multi-decadal dataset of daily river ice concentration (RIC) and annual phenology (freeze-up, breakup, and ice duration) for the six largest Arctic rivers—Yukon, Mackenzie, Ob, Yenisey, Lena, and Kolyma—covering hydrological years 2001–2024. Built from >590,000 MODIS Terra/Aqua scenes, our workflow integrates a scalable threshold-based classifier on Google Earth Engine with dual-satellite daily synthesis, temporal-window cloud reclassification, and a high-latitude dark-period correction. Technical validation against higher-resolution optical imagery shows a mean RIC accuracy of 0.83 across basins. Phenological metrics derived from MODIS agree with in situ records with mean absolute errors (MAE) of 10.8 days for freeze-up and 11.4 days for breakup (improving to 8.4 days relative to the onset of ice drift), and with Landsat-based river-section phenology with MAE of 10.5 days (freeze-up) and 16.0 days (breakup). RIC correlates strongly with surface air temperature (mean Pearson r = −0.91) and increases systematically with latitude. Trend analysis from 2001 through 2024 shows delayed freeze-up in over 66 % of river segments, earlier breakup in more than 65 %, and shorter ice seasons in over 65 %. On average, freeze-up is delayed by 9.0 days, breakup occurs 7.8 days earlier, and ice duration shortens by 14.1 days over the study period. These basin-consistent, temporally resolved records provide an open benchmark for diagnosing cryospheric change in Arctic river corridors and for constraining model–data intercomparisons. The river-ice dataset is available via Zenodo (https://doi.org/10.5281/zenodo.17054619, Qiu et al., 2025).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".