Are Terrestrial Biosphere Models Fit for Simulating the Global Land Carbon Sink?
Bibliographic record
Abstract
Abstract The Global Carbon Project estimates that the terrestrial biosphere has absorbed about one‐third of anthropogenic CO 2 emissions during the 1959–2019 period. This sink‐estimate is produced by an ensemble of terrestrial biosphere models and is consistent with the land uptake inferred from the residual of emissions and ocean uptake. The purpose of our study is to understand how well terrestrial biosphere models reproduce the processes that drive the terrestrial carbon sink. One challenge is to decide what level of agreement between model output and observation‐based reference data is adequate considering that reference data are prone to uncertainties. To define such a level of agreement, we compute benchmark scores that quantify the similarity between independently derived reference data sets using multiple statistical metrics. Models are considered to perform well if their model scores reach benchmark scores. Our results show that reference data can differ considerably, causing benchmark scores to be low. Model scores are often of similar magnitude as benchmark scores, implying that model performance is reasonable given how different reference data are. While model performance is encouraging, ample potential for improvements remains, including a reduction in a positive leaf area index bias, improved representations of processes that govern soil organic carbon in high latitudes, and an assessment of causes that drive the inter‐model spread of gross primary productivity in boreal regions and humid tropics. The success of future model development will increasingly depend on our capacity to reduce and account for observational uncertainties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.026 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".