Comparison of linear and mixed-effect regression models and a <i>k</i>-nearest neighbour approach for estimation of single-tree biomass
Bibliographic record
Abstract
Allometric biomass models for individual trees are typically specific to site conditions and species. They are often based on a low number of easily measured independent variables, such as diameter in breast height and tree height. A prevalence of small data sets and few study sites limit their application domain. One challenge in the context of the actual climate change discussion is to find more general approaches for reliable biomass estimation. Therefore, nonparametric approaches can be seen as an alternative to commonly used regression models. In this pilot study, we compare a nonparametric instance-based k-nearest neighbour (k-NN) approach to estimate single-tree biomass with predictions from linear mixed-effect regression models and subsidiary linear models using data sets of Norway spruce ( Picea abies (L.) Karst.) and Scots pine ( Pinus sylvestris L.) from the National Forest Inventory of Finland. For all trees, the predictor variables diameter at breast height and tree height are known. The data sets were split randomly into a modelling and a test subset for each species. The test subsets were not considered for the estimation of regression coefficients nor as training data for the k-NN imputation. The relative root mean square errors of linear mixed models and k-NN estimations are slightly lower than those of an ordinary least squares regression model. Relative prediction errors of the k-NN approach are 16.4% for spruce and 14.5% for pine. Errors of the linear mixed models are 17.4% for spruce and 15.0% for pine. Our results show that nonparametric methods are suitable in the context of single-tree biomass estimation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".