Improving maps of forest aboveground biomass: A combined approach using machine learning with a spatial statistical model
Bibliographic record
Abstract
Abstract. Aboveground biomass (AGB) estimates at the plot level plays a major part in connecting accurate single-tree AGB measurements to relatively difficult regional-scale AGB estimates. However, complex and spatially heterogeneous landscapes, where multiple environmental covariates (such as longitude, latitude, and forest structure) affect the spatial distribution of AGB, make upscaling of plot-level models more challenging. To address this challenge, this study proposes an approach that combines machine learning with spatial statistics to construct a more accurate plot-level AGB model. The study was conducted in a Eucalyptus plantation in Nanjing, China. We developed, evaluated, and compared the accuracy and performance of three different machine learning models [support vector machine (SVM), random forest (RF), and the radial basis function artificial neural network (RBF-ANN)], one spatial statistics model (P-BSHADE), and three combinations thereof (SVM & P-BSHADE, RF & P-BSHADE, RBF-ANN & P-BSHADE) for forest AGB estimates based on AGB data from 30 sample plots and their corresponding environmental covariates. The results show that the performance indices RMSE, nRMSE, MAE, and MRE of all combined models are substantially smaller than those of any individual models, with the RF & P-BSHADE combined method giving the smallest value. These results demonstrate clearly that combined models, especially the RF & P-BSHADE model, can improve the accuracy of plot-level AGB models and reduce uncertainty on plot-level AGB estimates or even on large-forested-landscape AGB estimates. These research results are important because they reduce the uncertainty in estimates of the regional carbon balance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".