Integrating remotely sensed data into forest resource inventories
Notice bibliographique
Résumé
The past two decades have demonstrated a great potential for airborne Light Detection and Ranging (LiDAR) data to improve the efficiency of forest resource inventories (FRIs). In order to make efficient use of LiDAR data in FRIs, the data need to be related to observations taken in the field. Various modeling techniques are available that enable a data analyst to establish a link between the two data sources. While the choice for a modeling technique may have negligible effects on point estimates, different model techniques may deliver different estimates of precision. This study investigated the impact of various model and variable selection procedures on estimates of precision. The focus was on LiDAR applications in FRIs. The procedures considered included stepwise variable selection procedures such as the Akaike Information Criterion (AIC), the corrected Akaike Information Criterion (AICc), and the Bayesian (or Schwarz) Information Criterion. Variables have also been selected based on the condition number of the matrix of covariates (i.e., LiDAR metrics) and the variance inflation factor. Other modeling techniques considered in this study were ridge regression, the least absolute shrinkage and selection operator (Lasso), partial least squares regression, and the random forest algorithm. Stepwise variable selection procedures have been considered in both, the (design-based) model-assisted, as well as in the model-based (or model-dependent) inference framework. All other techniques were investigated only for the model-assisted approach. In a comprehensive simulation study, the effects of the different modeling techniques on the precision of population parameter estimates (mean aboveground biomass per hectare) were investigated. Five different datasets were used. Three artificial datasets were simulated; two further datasets were based on FRI data from Canada and Norway. Canonical vine copulas were employed to create synthetic populations from the FRI data. From all populations simple random samples of different size were repeatedly drawn and the mean and variance of the mean were estimated for each sample. While for the model-based approach only a single variance estimator was investigated, for the model-assisted approach three alternative estimators were examined. The results of the simulation studies suggest that blind application of stepwise variable selection procedures lead to overly optimistic estimates of precision in LiDAR-assisted FRIs. The effects were severe for small sample sizes (n = 40 and n = 50). For large samples (n = 400) overestimation of precision was negligible. Good performance in terms of empirical standard errors and coverage rates were obtained for ridge regression, Lasso, and the random forest algorithm. This study concludes that the use of the latter three modeling techniques may prove useful in future LiDAR-assisted FRIs.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,012 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,003 | 0,007 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,003 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».