Assimilation of multiple datasets results in large differences in regional- to global-scale NEE and GPP budgets simulated by a terrestrial biosphere model
Bibliographic record
Abstract
Abstract. In spite of the importance of land ecosystems in offsetting carbon dioxide emissions released by anthropogenic activities into the atmosphere, the spatiotemporal dynamics of terrestrial carbon fluxes remain largely uncertain at regional to global scales. Over the past decade, data assimilation (DA) techniques have grown in importance for improving these fluxes simulated by terrestrial biosphere models (TBMs), by optimizing model parameter values while also pinpointing possible parameterization deficiencies. Although the joint assimilation of multiple data streams is expected to constrain a wider range of model processes, their actual benefits in terms of reduction in model uncertainty are still under-researched, also given the technical challenges. In this study, we investigated with a consistent DA framework and the ORCHIDEE-LMDz TBM–atmosphere model how the assimilation of different combinations of data streams may result in different regional to global carbon budgets. To do so, we performed comprehensive DA experiments where three datasets (in situ measurements of net carbon exchange and latent heat fluxes, spaceborne estimates of the normalized difference vegetation index, and atmospheric CO2 concentration data measured at stations) were assimilated alone or simultaneously. We thus evaluated their complementarity and usefulness to constrain net and gross C land fluxes. We found that a major challenge in improving the spatial distribution of the land C sinks and sources with atmospheric CO2 data relates to the correction of the soil carbon imbalance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".