ABCFlux v2: Arctic–boreal CO <sub>2</sub> and CH <sub>4</sub> monthly flux observations and ancillary information across terrestrial and freshwater ecosystems
Bibliographic record
Abstract
Abstract. Measurements of surface-atmosphere carbon dioxide (CO2) and methane (CH4) fluxes have been relatively sparse across the Arctic tundra and boreal biomes, causing significant uncertainties in carbon budget estimates from the region. While the availability of Arctic-boreal carbon flux data has increased substantially over the past decade, the data have remained spread across different repositories, scientific articles, and unpublished sources, making it difficult to leverage. Here we present a new dataset of monthly Arctic-boreal carbon fluxes (ABCFlux v2) across terrestrial (wetlands and uplands) and freshwater (lakes and rivers) ecosystems compiled from previous syntheses including the Arctic-boreal CO2 flux database (ABCFlux v1), the Boreal-Arctic Wetland and Lake Methane Dataset (BAWLD-CH4), and the Global River Methane Database (GRiMeDB). In addition, we consider data from general-purpose (e.g., Zenodo) and flux network repositories, literature, and site principal investigators. The dataset includes surface-atmosphere CO2 fluxes of gross primary production (GPP), ecosystem respiration (Reco), and net ecosystem exchange (NEE), alongside CH4 fluxes. For aquatic ecosystems, we split CH4 fluxes into diffusive and ebullitive flux pathways, and included potential emissions from transient storage in the water column (“storage fluxes”), alongside CO2 and CH4 concentrations dissolved in the surface water. Fluxes are measured through a variety of methods including chamber and eddy covariance techniques alongside bubble traps, ice-surveys, and concentration-based turbulence-driven modelling in aquatic ecosystems. The monthly flux data are reported together with supporting methodological and environmental metadata. The resulting ABCFlux v2 has 23,656 flux site-months, 8,182 concentration site-months, and 199 seasonal observations from 1,024 sites, and includes 55,560 reported fluxes (i.e. sum of GPP, Reco, NEE, and CH4 fluxes) from the years 1984 to 2024. The majority of monthly observations occurred after 1999. Wetlands had the highest number of site-month observations (8,641), followed by boreal forest (6,981), lotic ecosystems (6,275), lentic ecosystems (3,725) and upland tundra (3,308). Measurements of CO2 dominated the dataset across most ecosystem types (25,101) except for lentic ecosystems, where CH4 flux site-months (3,024) were more frequent than CO2 flux site-months (2,858). Overall, ABCFlux v2 includes 158 % more site-months for terrestrial CO2 flux data compared to ABCFlux v1. Integrating and updating BAWLD-CH4 flux data from growing season averages to monthly fluxes resulted in 5,671 site-months of chamber CH4 data compared to 762 site-years. This collaborative initiative, involving contributions from over 260 researchers, provides a comprehensive overview of the current state of the Arctic-boreal carbon flux network and its data, and serves as an important step in reducing uncertainties in Arctic-boreal carbon budgets and in enhancing our understanding of climate feedbacks. The data can be accessed at ORNL DAAC at https://doi.org/10.3334/ORNLDAAC/2448 (Virkkala et al., 2025b).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".