Retrieving Inland Water Quality Parameters via Satellite Remote Sensing: Sensor Evaluation, Atmospheric Correction, and Machine Learning Approaches
Bibliographic record
Abstract
Satellite remote sensing provides a cost-effective and large-scale alternative to traditional methods for retrieving water quality parameters for inland waters. Effective water quality parameter retrieval via optical satellite remote sensing requires three key components: (1) a sensor whose measurements are sensitive to variations in water quality; (2) accurate atmospheric correction to eliminate the effect of absorption and scattering in the atmosphere and retrieve the water-leaving radiance/reflectance; and (3) a bio-optical model used to estimate water quality from the optical signal. This study provides a literature review and an evaluation of these three components. First, a review of decommissioned, active, and upcoming satellite sensors is presented, highlighting their advantages and limitations, and a ranking method is introduced to assess their suitability for retrieving chlorophyll-a, colored dissolved organic matter, and non-algal particles in inland waters. This ranking can aid in selecting appropriate sensors for future studies. Second, the strengths and weaknesses of atmospheric correction algorithms used over inland waters are examined. The results show that no atmospheric correction algorithm performed consistently across all conditions. However, understanding their strengths and weaknesses allows users to select the most suitable algorithm for a specific use case. Third, the challenges, limitations, and recent advances of machine learning use in bio-optical models for inland water quality parameter retrieval are discussed. Machine learning models have limitations, including low generalizability, low dimensionality, spatial/temporal autocorrelation, and information leakage. These issues highlight the importance of locally trained models, rigorous cross-validation methods, and integrating auxiliary data to enhance dimensionality. Finally, recommendations for promising research directions are provided.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".