MétaCan
Menu
Back to cohort
Record W4384341039 · doi:10.1111/2041-210x.14167

A deep learning approach to photo–identification demonstrates high performance on two dozen cetacean species

2023· article· en· W4384341039 on OpenAlexaff
Philip T. Patton, Ted Cheeseman, Kenshin Abe, Taiki Yamaguchi, Walter Reade, Ken Southerland, Addison Howard, Erin M. Oleson, Jason B. Allen, Erin Ashe, Aline Athayde, Robin W. Baird, Charla J. Basran, Elsa Cabrera, John Calambokidis, Júlio Cardoso, Emma L. Carroll, Amina Cesario, Barbara Cheney, Enrico Corsi, Jens J. Currie, John W. Durban, Erin A. Falcone, Holly Fearnbach, Kiirsten Flynn, Trish Franklin, Wally Franklin, Bárbara Galletti Vernazzani, Tilen Genov, Marie C. Hill, David Johnston, Erin L. Keene, Sabre D. Mahaffy, Tamara L. McGuire, Liah McPherson, Catherine Meyer, Robert Michaud, Anastasia Miliou, Dara N. Orbach, Heidi C. Pearson, Marianne H. Rasmussen, William Rayment, Caroline Rinaldi, Renato Rinaldi, Salvatore Siciliano, Stephanie H. Stack, Beatriz Tintoré, Leigh G. Torres, Jared R. Towers, Cameron Trotter, Reny B. Tyson, Caroline R. Weir, Rebecca Wellard, Randall S. Wells, Kymberly M. Yano, Jochen R. Zaeschmar, Lars Bejder

Bibliographic record

VenueMethods in Ecology and Evolution · 2023
Typearticle
Languageen
FieldEnvironmental Science
TopicMarine animal studies overview
Canadian institutionsBayer (Canada)
FundersNational Marine Fisheries ServiceNational Oceanic and Atmospheric AdministrationNational Science Foundation
KeywordsDozenIdentification (biology)Deep learningBiologyEvolutionary biologyArtificial intelligenceGeographyComputer scienceEcologyMathematics

Abstract

fetched live from OpenAlex

Abstract Researchers can investigate many aspects of animal ecology through noninvasive photo–identification. Photo–identification is becoming more efficient as matching individuals between photos is increasingly automated. However, the convolutional neural network models that have facilitated this change need many training images to generalize well. As a result, they have often been developed for individual species that meet this threshold. These single‐species methods might underperform, as they ignore potential similarities in identifying characteristics and the photo–identification process among species. In this paper, we introduce a multi‐species photo–identification model based on a state‐of‐the‐art method in human facial recognition, the ArcFace classification head. Our model uses two such heads to jointly classify species and identities, allowing species to share information and parameters within the network. As a demonstration, we trained this model with 50,796 images from 39 catalogues of 24 cetacean species, evaluating its predictive performance on 21,192 test images from the same catalogues. We further evaluated its predictive performance with two external catalogues entirely composed of identities that the model did not see during training. The model achieved a mean average precision (MAP) of 0.869 on the test set. Of these, 10 catalogues representing seven species achieved a MAP score over 0.95. For some species, there was notable variation in performance among catalogues, largely explained by variation in photo quality. Finally, the model appeared to generalize well, with the two external catalogues scoring similarly to their species' counterparts in the larger test set. From our cetacean application, we provide a list of recommendations for potential users of this model, focusing on those with cetacean photo–identification catalogues. For example, users with high quality images of animals identified by dorsal nicks and notches should expect near optimal performance. Users can expect decreasing performance for catalogues with higher proportions of indistinct individuals or poor quality photos. Finally, we note that this model is currently freely available as code in a GitHub repository and as a graphical user interface, with additional functionality for collaborative data management, via Happywhale.com.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.038
Threshold uncertainty score0.404

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.033
GPT teacher head0.310
Teacher spread0.276 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations27
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueMethods in Ecology and EvolutionSame topicMarine animal studies overviewFrench-language works237,207