Continental-scale mapping of soil pH with SAR-optical fusion based on long-term earth observation data in google earth engine
Bibliographic record
Abstract
• The selection of satellite sensor and radar system parameters greatly affected the model output. • The model was improved when more polarizations, orbital directions and frequencies were involved. • Models built using multiband radar datasets achieved a comparable accuracy to models based on optical data. • The model built by fusing SAR-optical data achieved better results than the model without SAR data. • The development of this continental-scale DSM work largely benefits from GEE. The advent of cloud computing platforms (e.g., Google Earth Engine (GEE)) and the massive amounts of optical and radar Earth Observation (EO) data hosted by these platforms present new opportunities for mapping soil pH at large scales. However, existing studies generally lack consensus on the effects of satellite and radar sensor parameters on GEE-based soil pH prediction models. In this study, we assessed the suitability of long-term radar (C-band Sentinel-1 and L-band PALSAR-1/2) and optical (Sentinel-2) EO data on GEE for the digital mapping of soil pH on a continental (Europe) scale and determined the most appropriate radar sensor parameters. Thirteen scenarios with different data configurations were simulated and combined with the 2018 LUCAS soil database and two machine learners (boosted regression trees and extreme gradient boosting) to develop soil prediction models. Results showed that the selection of modeling techniques, satellite sensors and radar system parameters largely affected the model output. Models involving a single polarization mode of PALSAR-1/2 data performed the worst (RPD = 1.24). Models based on Sentinel-1 data performed better than those built using PALSAR-1/2 data. The model performance was improved when a model involved more polarization bands, orbital directions, and band frequencies. The multiband model built using the two radar datasets achieved a comparable accuracy to the model based on optical data. Moreover, the model that fused radar-optical data achieved better results, with RPD values of 1.56 and 1.46 for the models with and without radar data, respectively; its performance was comparable to that of models built with commonly used variables (topography and climate). The analysis of importance indicated that long-term optical and radar EO data on GEE were important in our model. The modelling of soil pH at the continental scale largely benefits from GEE. The predicted maps exhibited strong spatial heterogeneity among different biogeographic regions, with similar spatial patterns under different modelling scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".