Predicting dreissenid mussel abundance in nearshore waters using underwater imagery and deep learning
Bibliographic record
Abstract
Abstract Accurate and cost‐effective dreissenid mussel abundance maps are vital to assess their ecological roles in aquatic systems. A deep neural network (DNN) modeling framework using semantic segmentation was developed to automatically assess the abundance distribution of two invasive mussel species: zebra and quagga. DNN models were trained on images captured in Lake Erie and Lake Ontario using an underwater color imaging technique. The accuracy of the method was assessed relative to manual laboratory counts of harvested mussels, their dry biomass, and percentage live coverage estimated from fixed‐size quadrats. Assessments performed on a test set collected from 2016 to 2018 show that DNN‐based mussel coverage predictions explain 79% of the variance in log biomass, and 71% for log abundance ( N = 125). For reference, live coverage estimated by scuba divers was transformed and found to be a better predictor of biomass (93%) and abundance (91%) ( N = 725), leaving room for improvement of our automated method. When identical images were presented to eight human analysts and the DNN, the agreement in live mussel coverage prediction was 85% ( N = 189). Models generalize well to diverse underwater illuminations, camera orientations, and resolutions, but are adversely impacted by occluding vegetation and suspended sediment. DNN models are an efficient and accurate solution for mapping mussel abundances at a scale that was previously impossible. The method may be integrated with other studies to assess the mussels' impacts in a variety of aquatic ecosystems. Source code: https://github.com/AngusG/deep-learning-dreissenid and data https://doi.org/10.5683/SP3/MZEBOJ for reproducing our method are publicly available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".