A Comparative Analysis of Machine Learning Algorithms for Tree Species Recognition Using An Image-Based Approach with Implementation Potential for Close-range Technologies
Notice bibliographique
Résumé
Close-range technologies capable of capturing forest ecosystems in three-dimensional space with great detail are revolutionising precision forestry research and practice, mainly by increasing the level of automation for data collection and processing. Furthermore, they provide options to measure some parameters directly, for example, volume or biomass. However, automatic tree species recognition still needs to be properly solved, which is a crucial and challenging task. A couple of approaches by different authors were done to overcome the challenge when data from close-range technologies are used. The authors mainly utilised 3D structures of whole trees or, in some cases, bark structures using point clouds. Or derived 2D blueprints of whole trees from point clouds to distinguish between tree species. In our approach, we are using images of bark. Usually, images are taken during the data acquisition by close-range technologies as a resource for photogrammetry or for colourising the point clouds in the case of terrestrial laser scanning, for example. Carpentier et al. (2018) did an experiment with 23 tree species in Canada and used convolutional neural networks to classify tree species with an accuracy of almost 94%. We focused on benchmarking multiple machine learning and deep learning algorithms in our experiment. Namely: Random forest; Decision tree; Support Vector Machine; Gradient boost; K-nearest Neighbors; Gaussian Naïve Bayes; Multilayer Perceptron; Convolutional neural networks.In our first experiment, we collected two datasets of bark images using Sony alfa 7 and Canon EOS 4000D. We have collected 1755 images in Slovakia (1369) and Czechia (386); both datasets contain four tree species. The four species from Slovak datasets are European beech, sessile oak, Norway spruce, and European silver fir. Czechia data consists of the species European beech, large-leaved linden, Norway maple, and Scots pine. However, the bark images from Slovakia are from managed forests, and there is a variety of markings on bark; for that, images are cropped to small regions excluding the markings.The most accurate results were achieved by CNN, which provides 94% accuracy on Slovak exact cropped dataset with a 50% dropout and 91% on an exact cropped dataset with a 50% dropout. When CNN is not considered, the most accurate algorithm was Multilayer perceptron with an accuracy of 92%.The following research will focus on implementing such tree species classification within the point cloud processing workflow when close-range technologies are used. Secondly, Carpentier et al. (2018) created Barknet 1.0, where they stored 23,000 high-resolution bark images of 23 tree species in Canada. Our next goal is to develop a database of tree species across Europe. To achieve such a challenging task, we will do it within the 3DForEcoTech COST Action, a European collaborative project focusing on close-range technologies and their implementation for precision forestry and forest ecology.ReferencesCarpentier, M., Giguere, P. and Gaudreault, J., 2018, October. Tree species identification from bark images using convolutional neural networks. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 1075-1081). IEEE.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,012 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,004 | 0,003 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,002 | 0,003 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».