Evaluating the agreement between measurements and models of net ecosystem exchange at different times and time scales using wavelet coherence: an example using data from the North American Carbon Program Site-Level Interim Synthesis
Bibliographic record
Abstract
Abstract. Earth system processes exhibit complex patterns across time, as do the models that seek to replicate these processes. Model output may or may not be significantly related to observations at different times and on different frequencies. Conventional model diagnostics provide an aggregate view of model-data agreement, but usually do not identify the time and frequency patterns of model misfit, leaving unclear the steps required to improve model response to environmental drivers that vary on characteristic frequencies. Wavelet coherence can quantify the times and frequencies at which models and measurements are significantly different. We applied wavelet coherence to interpret the predictions of twenty ecosystem models from the North American Carbon Program (NACP) Site-Level Interim Synthesis when confronted with eddy covariance-measured net ecosystem exchange (NEE) from ten ecosystems with multiple years of available data. Models were grouped into classes with similar approaches for incorporating phenology, the calculation of NEE, and the inclusion of foliar nitrogen (N). Models with prescribed, rather than prognostic, phenology often fit NEE observations better on annual to interannual time scales in grassland, wetland and agricultural ecosystems. Models that calculate NEE as net primary productivity (NPP) minus heterotrophic respiration (HR) rather than gross ecosystem productivity (GPP) minus ecosystem respiration (ER) fit better on annual time scales in grassland and wetland ecosystems, but models that calculate NEE as GPP – ER were superior on monthly to seasonal time scales in two coniferous forests. Models that incorporated foliar nitrogen (N) data were successful at capturing NEE variability on interannual (multiple year) time scales at Howland Forest, Maine. Combined with previous findings, our results suggest that the mechanisms driving daily and annual NEE variability tend to be correctly simulated, but the magnitude of these fluxes is often erroneous, suggesting that model parameterization must be improved. Few NACP models correctly predicted fluxes on seasonal and interannual time scales where spectral energy in NEE observations tends to be low, but where phenological events, multi-year oscillations in climatological drivers, and ecosystem succession are known to be important for determining ecosystem function. Mechanistic improvements to models must be made to replicate observed NEE variability on seasonal and interannual time scales.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".