Bulk arthropod abundance, biomass and diversity estimation using deep learning for computer vision
Bibliographic record
Abstract
Abstract Arthropod abundance, biomass and taxonomic diversity are key metrics often used to assess the efficacy of restoration efforts. Gathering these metrics is a slow and laborious process, quantified by an expert manually sorting and weighing arthropod specimens. We present a tool to accelerate bulk arthropod classification and biomass estimates utilizing machine learning methods for computer vision. Our approach requires pre‐sorted arthropod samples to create a training dataset. We construct a dataset considering 18 terrestrial arthropod functional groups collected in southern Ontario, Canada. The dataset contains 517 high‐resolution images with approximately 20 individuals per image taken from either a petri dish or a bulk tray. Our tool uses the watershed algorithm to obtain precisely cropped individuals without any object annotations. After manually sorting cropped images of biological ‘debris’ and petri dish edges, three classifiers, DenseNet121, ResNet101 and MobileNetv2, each with trade‐offs of computational efficiency versus accuracy, are trained and compared to predict arthropod functional groups for each cropped individual. To calculate biomass, we compare seven linear and nonlinear models considering the arthropod pixel masks obtained using the watershed algorithm, in combination with images of a single function group with recorded weights, to calculate the per pixel density per functional group. From our experimentation, we recommend using DenseNet121 as it had the highest top‐1 functional group classification accuracy, likely a result of being the model with the largest number of parameters, with 86.14% considering the 20 labelled classes (18 arthropods plus debris and petri dish edge) in comparison to ResNet101 (85.10%) and MobileNetv2 (84.94%). For biomass estimation, we recommend using the average per pixel density which had the highest ranked performance considering both total error, 0.043 g (0.855% error), and cumulative class‐specific error, 1.62 g (40.67% average error across all classes), in comparison to the total ground truth biomass of 5.10 g. Our estimated Simpson's Index of Diversity was 0.9404 in comparison to the ground truth 0.9408. Our method simultaneously classifies >1,000 arthropods to functional groupings while estimating total and class specific biomass, without any computer vision bounding box or mask annotations, all from a single photo. We release our code and dataset to further research efforts in computer vision for arthropod classification.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".