Automated surface water extraction from RapidEye imagery including cloud and cloud shadow detection
Bibliographic record
Abstract
Mapped surface water extents represent fundamental geospatial information required by a large number of public and private sector stakeholders in Canada. Currently available surface water maps from Canada's National Hydrographic Network (NHN) are out-of-date in many locations in part due to the vintage of the maps themselves, and also because water extents are dynamic and changing. Previously, a significant amount of Canada's NHN base geospatial data was manually interpreted from airphotos and other sources of high-resolution imagery, which enabled detailed mapping but required significant human and financial resources. Satellite imagery can provide a cheaper alternative due to the free availability of certain data and potential for automated feature extraction. Freely available medium resolution data from sensors such as Landsat can provide information at a scale of 1:50 k while more detail is needed to map smaller water courses that can only be resolved from high resolution sensors (< 5 m) such as RapidEye. This Open File describes a robust, fully automated procedure to extract surface water from RapidEye imagery including cloud and cloud shadow screening, as a potential method to enhance future NHN updating. The method is applied to RapidEye imagery residing in the Government of Canada's image archive over the St-John, Red and Richelieu Rivers. Qualitative assessment of the Richelieu product shows high overall agreement with water features in Google Earth imagery and NHN, with errors stemming from spectral confusion with dark soil and urban shadow. Canada's RapidEye archive currently covers approximately 30 % of provinces while its 5 m spatial resolution provides a good compromise between detail and data volume that can be processed on modern workstations. Similar spectral bands available in other higher resolution satellite sensors (< 2 m) such as WorldView suggest more detailed surface water extents may become available using methods developed here and adapted to these sensors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.010 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".