Estimating time-dependent vegetation biases in the SMAP soil moisture product
Bibliographic record
Abstract
Abstract. Remotely sensed soil moisture products are influenced by vegetation and how it is accounted for in the retrieval, which is a potential source of time-variable biases. To estimate such complex, time-variable error structures from noisy data, we introduce a Bayesian extension to triple collocation in which the systematic errors and noise terms are not constant but vary with explanatory variables. We apply the technique to the Soil Moisture Active Passive (SMAP) soil moisture product over croplands, hypothesizing that errors in the vegetation correction during the retrieval leave a characteristic fingerprint in the soil moisture time series. We find that time-variable offsets and sensitivities are commonly associated with an imperfect vegetation correction. Especially the changes in sensitivity can be large, with seasonal variations of up to 40 %. Variations of this size impede the seasonal comparison of soil moisture dynamics and the detection of extreme events. Also, estimates of vegetation–hydrology coupling can be distorted, as the SMAP soil moisture has larger R2 values with a biomass proxy than the in situ data, whereas noise alone would induce the opposite effect. This observation highlights that time-variable biases can easily give rise to distorted results and misleading interpretations. They should hence be accounted for in observational and modelling studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".