Building a Land Data Assimilation Community to Tackle Technical Challenges in Quantifying and Reducing Uncertainty in Land Model Predictions
Bibliographic record
Abstract
he regional to global-scale process-based land models that form part of numerical weather prediction (NWP) systems or Earth system models (ESMs) have rapidly increased in their complexity over the past few decades (Prentice et al. 2015;Fisher and Koven, 2020).In parallel, there is a growing wealth of terrestrial observations that can be used to confront land models, including long-running observation/experiment campaigns [e.g., NSF Long-Term Ecological Research (LTER) sites and DOE Next-Generation Ecosystem Experiments-Tropics (NGEE-Tropics) and Arctic] and satellite missions or products [e.g., Landsat and MODIS from NASA, ESA Climate Change Initiative (CCI) Soil Moisture], the synthesis of site-based and experimental manipulation data into networks [e.g., FLUXNET, Integrated Carbon Observation System (ICOS), the International Soil Moisture Network, SAPFLUXNET, Drought-Net, Free Air CO 2 Enrichment (FACE)], novel ground-based observations (e.g., tree ring data, carbonyl sulfide, radiocarbon measurements), and a new array of space-based observations of global carbon, water, and energy cycles [e.g., hyperspectral and lidar instruments such as Global Ecosystem Dynamics Investigation (GEDI) and Hyperspectral Imager Suite (HISUI) on the International Space Station (ISS), and solar induced fluorescence from platforms like the Orbiting Carbon Observatory 2 (OCO-2) and Tropospheric Monitoring Instrument (TROPOMI)] at higher spatial, temporal, and spectral resolutions than ever before.Despite the increasing complexity of models and density of land-based data, uncertainty in land model projections remains high.While there has been a concerted effort to use terrestrial observations for model evaluation, the number of studies that use these data for quantifying and reducing uncertainty in land model parameters and states via a statistical data assimilation (DA) framework is small in comparison.This is primarily due to the computational expense and technical challenges associated with implementing such a global-scale land DA system.However, land models urgently need to be confronted with a wide range of data to optimize model parameters, initialize surface states, and to address model structural uncertainty.Without such efforts, we cannot quantify or reduce uncertainty associated with individual model projections, and the intermodel spread in weather forecasts and predictions of land-atmospheric interactions or carbon-climate feedbacks will remain high (Arora et al. 2020). The challenge of developing land DA systemsA number of land modeling groups spanning different modeling communities [carbon, hydrology, land surface modeling (LSM)/ESM, and NWP] have devoted significant resources into developing global-scale land DA systems.However, this technical development work does not get the level of exposure in the literature or in conference talks commensurate with the resources needed to complete that work because publications and presentations are naturally focused on scientific questions.For the same reason, development of DA systems is typically not the focus of grant proposals nor calls for proposals by funding agencies.Nonetheless
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.012 |
| Open science | 0.004 | 0.009 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".