Cross-site comparison of ecosystem- and plot-scale methane fluxes from wetlands and uplands (Version 1)
Bibliographic record
Abstract
This dataset contains the six temporal aggregations (separate datasets; half-hourly, hourly, daily, weekly, monthly and annual) used in the article "Määttä et al. A cross-site comparison of ecosystem- and plot-scale methane fluxes from wetlands and uplands" (manuscript submitted to Biogeosciences 2025-10-10). The datasets contain contemporaneous timestamp-aligned methane (CH4) flux data from eddy covariance (EC) and chamber measurements from 10 sites, as well as environmental data derived from FLUXNET-CH4 (Delwiche et al., 2020; Knox et al., 2019; https://fluxnet.org/data/fluxnet-ch4-community-product/) and chamber-associated measurements used as predictors of ecosystem and plot-scale CH4 flux differences. Please note that all (except for site US-StJ) EC data were obtained from the FLUXNET-CH4 database (Delwiche et al., 2020; Knox et al., 2019; CC-BY 4.0) and chamber CH4 flux and environmental data from site data providers (chamber data that has not been published elsewhere with CC-BY 4.0 has been included in these datasets with permission from data providers). EC data for US-StJ were obtained from the site data providers. For questions related to site-specific chamber and EC CH4 flux and environmental data, please contact the data providers or refer to site-specific literature (see README.txt).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".