A map of global peatland distribution created using machinelearning for use in terrestrial ecosystem and earth system models
Bibliographic record
Abstract
Abstract. Peatlands store large amounts of soil carbon and constitute an important component of the global carbon cycle. Accurate information on the global extent and distribution of peatlands is presently lacking but it important for earth system models (ESMs) to be able to simulate the effects of climate change on the global carbon balance. The most comprehensive peatland map produced to date is a qualitative presence/absence product. Here, we present a spatially continuous global map of peatland fractional coverage using the extremely randomized tree machine learning method suitable for use as a prescribed geophysical field in an ESM. Inputs to our statistical model include spatially distributed climate data, soil data and topographical slopes. Available maps of peatland fractional coverage for Canada and West Siberia were used along with a proxy for non-peatland areas to train and test the statistical model. Regions where the peatland fraction is expected to be zero were estimated from a map of topsoil organic carbon content below a threshold value of 13 kg/m2. The modelled coverage of peatlands yields a root mean square error of 4 % and a coefficient of determination of 0.91 for the 10,978 tested 0.5 degree grid cells. We then generated a complete global peatland fractional coverage map. In comparison with earlier qualitative estimates, our global modelled peatland map is able to reproduce peatland distributions in places remote from the training areas and capture peatland hot spots in both boreal and tropical regions, as well as in the southern hemisphere. Additionally we demonstrate that our machine-learning method has greater skill than solely setting peatland areas based on histosols from a soil database.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".