MétaCan
Menu
Back to cohort

The PanAf-FGBG Dataset

2025· dataset· W7135417822 on OpenAlexaboutno aff
Otto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens, Paula Dieguez, Thurston C. Hicks, Sorrel Jones, Kevin Lee, Maureen S. McCarthy, Amelia Meier, Emmanuelle Normand, Erin G. Wessling, Roman M.Wittig, Kevin E. Langergraber, Klaus Zuberbuhler, Lukas Boesch, Thomas Schmid, Mimi Arandjelovic, Hjalmar S. Kühl, Tilo Burghardt

Bibliographic record

VenueBristol Research (University of Bristol) · 2025
Typedataset
Language
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsMetadataCitationEndangered speciesWildlifeResource (disambiguation)Geocoding

Abstract

fetched live from OpenAlex

DESCRIPTION. The PanAf-FGBG dataset comprises behaviour-annotated video footage of wild chimpanzees from more than 350 camera locations across tropical Africa, collected by the Pan African Programme: The Cultured Chimpanzee. It includes paired foreground (with chimpanzees) and background (without chimpanzees) videos, allowing controlled analysis of background influence on behaviour recognition models. The dataset is split into overlapping and disjoint camera location views to support evaluation under both in-distribution and out-of-distribution conditions. Each entry is accompanied by metadata and multi-label annotations for 14 distinct behaviours, enabling robust model training and testing. This resource aims to enhance AI models for wildlife behaviour understanding and supports broader conservation efforts for endangered great ape species. CITATION. When using this data please cite this dataset deposit and the associated paper where the dataset and baselines are explained in detail: "The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recognition" published in the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) available here: https://openaccess.thecvf.com/content/CVPR2025/papers/Brookes_The_PanAf-FGBG_Dataset_Understanding_the_Impact_of_Backgrounds_in_Wildlife_CVPR_2025_paper.pdf. For BIBTEX citation details please see the project website: https://obrookes.github.io/panaf-fgbg.github.io/ ACKNOWLEDGEMENTS. We thank the Pan African Programme: 'The Cultured Chimpanzee' team and its collaborators for allowing the use of their data for this paper. We thank Amelie Pettrich, Antonio Buzharevski, Eva Martinez Garcia, Ivana Kirchmair, Sebastian Schütte, Linda Gerlach and Fabina Haas. We also thank management and support staff across all sites; specifically Yasmin Moebius, Geoffrey Muhanguzi, Martha Robbins, Henk Eshuis, Sergio Marrocoli and John Hart. Thanks to the team at https://www.chimpandsee.org particularly Briana Harder, Anja Landsmann, Laura K. Lynn, Zuzana Macháčková, Heidi Pfund, Kristeena Sigler and Jane Widness. The work that allowed for the collection of the dataset was funded by the Max Planck Society, Max Planck Society Innovation Fund, and Heinz L. Krekeler. In this respect we would like to thank: Ministre des Eaux et Forêts, Ministère de l'Enseignement supérieur et de la Recherche scientifique in Côte d'Ivoire; Institut Congolais pour la Conservation de la Nature, Ministère de la Recherche Scientifique in Democratic Republic of Congo; Forestry Development Authority in Liberia; Direction Des Eaux Et Forêts, Chasses Et Conservation Des Sols in Senegal; Makerere University Biological Field Station, Uganda National Council for Science and Technology, Uganda Wildlife Authority, National Forestry Authority in Uganda; National Institute for Forestry Development and Protected Area Management, Ministry of Agriculture and Forests, Ministry of Fisheries and Environment in Equatorial Guinea. This work was supported by the UKRI CDT in Interactive AI (grant EP/S022937/1). This work was in part supported by the US National Science Foundation Awards No. 2118240 "HDR Institute: Imageomics: A New Frontier of Biological Information Powered by Knowledge-Guided Machine Learning" and Award No. 2330423 and Natural Sciences and Engineering Research Council of Canada under Award No. 585136 for the "AI and Biodiversity Change (ABC) Global Center". WEBSITE. Further materials are available at the project website at: https://obrookes.github.io/panaf-fgbg.github.io/

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.051
Threshold uncertainty score0.103

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.005
Meta-epidemiology (narrow)0.0040.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0040.005
Science and technology studies0.0020.001
Scholarly communication0.0030.003
Open science0.0040.002
Research integrity0.0040.002
Insufficient payload (model declined to judge)0.0310.060

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.075
GPT teacher head0.370
Teacher spread0.295 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBristol Research (University of Bristol)French-language works237,207