Bulk arthropod abundance, biomass and diversity estimation using deep learning for computer vision
Notice bibliographique
Résumé
Abstract Arthropod abundance, biomass and taxonomic diversity are key metrics often used to assess the efficacy of restoration efforts. Gathering these metrics is a slow and laborious process, quantified by an expert manually sorting and weighing arthropod specimens. We present a tool to accelerate bulk arthropod classification and biomass estimates utilizing machine learning methods for computer vision. Our approach requires pre‐sorted arthropod samples to create a training dataset. We construct a dataset considering 18 terrestrial arthropod functional groups collected in southern Ontario, Canada. The dataset contains 517 high‐resolution images with approximately 20 individuals per image taken from either a petri dish or a bulk tray. Our tool uses the watershed algorithm to obtain precisely cropped individuals without any object annotations. After manually sorting cropped images of biological ‘debris’ and petri dish edges, three classifiers, DenseNet121, ResNet101 and MobileNetv2, each with trade‐offs of computational efficiency versus accuracy, are trained and compared to predict arthropod functional groups for each cropped individual. To calculate biomass, we compare seven linear and nonlinear models considering the arthropod pixel masks obtained using the watershed algorithm, in combination with images of a single function group with recorded weights, to calculate the per pixel density per functional group. From our experimentation, we recommend using DenseNet121 as it had the highest top‐1 functional group classification accuracy, likely a result of being the model with the largest number of parameters, with 86.14% considering the 20 labelled classes (18 arthropods plus debris and petri dish edge) in comparison to ResNet101 (85.10%) and MobileNetv2 (84.94%). For biomass estimation, we recommend using the average per pixel density which had the highest ranked performance considering both total error, 0.043 g (0.855% error), and cumulative class‐specific error, 1.62 g (40.67% average error across all classes), in comparison to the total ground truth biomass of 5.10 g. Our estimated Simpson's Index of Diversity was 0.9404 in comparison to the ground truth 0.9408. Our method simultaneously classifies >1,000 arthropods to functional groupings while estimating total and class specific biomass, without any computer vision bounding box or mask annotations, all from a single photo. We release our code and dataset to further research efforts in computer vision for arthropod classification.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».