Assessment of a probabilistic supervised machine learning method to estimate biomass expansion and conversion factors: a case study on cedar and pine trees
Bibliographic record
Abstract
Quantifying tree and forest biomass is crucial for formulating effective forest policy and management, given its role in human resource use and carbon storage. Forest biomass significantly contributes to environmental quality by absorbing carbon dioxide. Current research focuses on accurately determining biomass factors for various tree species, reflecting the emphasis on estimating and predicting tree biomass and carbon stocks. This study employed both standard nonlinear regression modeling ( NLR) and Gaussian process regression ( GPR), a machine learning method using artificial intelligence, to estimate and predict biomass expansion and conversion factors accurately. The case study included plantation forests and naturally occurring cedar and pine trees in Türkiye’s Western Inner Anatolian Region and Göller Region (Northern Mediterranean Region). Nonlinear regression used the Levenberg-Marquardt optimization method, while GPR employed the radial basis function kernel. This dual approach allowed for assessing prediction uncertainties. The models constructed using GPR show superior performance compared to the NLR models for both biomass factors and species within the datasets used. According to the Furnival evaluation metric values, the accuracy of the NLR models was 1.05 to 1.34 times lower than that of the corresponding GPR models. The overall findings highlight the significant potential of GPR for accurately estimating and predicting biomass factors with high variances. This emphasizes its utility in modeling scenarios that require high flexibility, such as tree biomass prediction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".