MétaCan
Menu
Back to cohort
Record W4410836337 · doi:10.5194/jm-44-145-2025

Community guidelines to increase the reusability of marine microfossil assemblage data

2025· article· en· W4410836337 on OpenAlexaff
Lukas Jonkers, Anne Strack, Montserrat Alonso‐García, Simon D’haenens, Robert Huber, Michal Kučera, Iván Hernández‐Almeida, Chloe Jones, Brett Metcalfe, Rajeev Saraswat, Lóránd Silye, Sanjay Kumar Verma, Muhamad Naim Abd Malek, Gerald Auer, Catia F Barbosa, María Ángeles Bárcena, Karl‐Heinz Baumann, Flavia Boscolo‐Galazzo, Joeven Austine S. Calvelo, Lucilla Capotondi, Martina Caratelli, Jorge Cardich, Humberto Carvajal-Chitty, Markéta Chroustová, Helen K. Coxall, Renata M. de Mello, Anne de Vernal, Paula Diz, Kirsty M. Edgar, Helena L. Filipsson, Ángela Fraguas, Heather Furlong, Giacomo Galli, Natalia García Chapori, Robyn Granger, Jeroen Groeneveld, Adil Imam, Rebecca Jackson, David Lazarus, Julie Meilland, Marína Molčan Matejová, Raphaël Morard, Caterina Morigi, Sven N. Nielsen, Diana Ochoa, Maria Rose Petrizzo, Andrés S. Rigual‐Hernández, Marina C. Rillo, Matthew L. Staitis, Gamze Tanık, Raúl Tapia, Nishant Vats, Bridget S. Wade, Alexander Weinmann

Bibliographic record

VenueJournal of Micropalaeontology · 2025
Typearticle
Languageen
FieldEnvironmental Science
TopicIsotope Analysis in Ecology
Canadian institutionsUniversité du Québec à Montréal
FundersNederlandse Organisatie voor Wetenschappelijk OnderzoekDeutsche Forschungsgemeinschaft
KeywordsReusabilityAssemblage (archaeology)OceanographyGeologyEnvironmental resource managementEnvironmental scienceComputer sciencePaleontology

Abstract

fetched live from OpenAlex

Data on marine microfossil assemblage composition have multiple applications. Initially, they were primarily used for (chrono)stratigraphy and palaeoecology, but these data are now also widely used to study evolutionary and ecological processes, such as past biodiversity and its links with environmental dynamics, or to provide a basis for conservation efforts and biomonitoring. The large range of potential applications renders microfossil abundance data ideal for reuse. However, the complexity inherent in taxonomic data, which encompass extant and extinct species, coupled with the inherent intricacies of information on biological communities extracted from sedimentary archives, poses considerable hurdles in reusing marine microfossil data, even when they are publicly available. Here, we present guidelines derived from an online survey conducted within the marine micropalaeontological community, aimed at improving the reusability of microfossil assemblage data. These guidelines advocate for clarity and transparency in the documentation of the methods and the outcome, and we outline the data attributes required for effective reuse of micropalaeontological data. These guidelines are intended for researchers who generate microfossil abundance datasets and for reviewers, editors, and data curators at repositories. A total of 113 researchers evaluated the relevance of about 50 data attributes that might be needed to enable and maximise the reuse of marine microfossil abundance datasets. Each property is ranked based on the survey results. All information is, in principle, considered “desired”. Information that improves the reusability is ranked as “recommended”, and information that is required for reuse is ranked as “essential”. Analysis of a selection of datasets available online reveals a rather large gap between data properties deemed essential by survey participants and what is actually contained in publicly available microfossil assemblage datasets. While the survey indicates that the micropalaeontological community values good data stewardship, improving data reusability still requires new efforts to incorporate all the essential information. The guidelines presented here are intended as a step in that direction. Determining the optimal forms and formats for data sharing are obvious next steps the community needs to take.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.441
metaresearch head score (Gemma)0.735
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.988
Threshold uncertainty score0.690

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4410.735
Meta-epidemiology (narrow)0.0020.005
Meta-epidemiology (broad)0.0030.005
Bibliometrics0.0370.024
Science and technology studies0.0070.008
Scholarly communication0.0140.018
Open science0.0120.019
Research integrity0.0120.010
Insufficient payload (model declined to judge)0.0090.011

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.043
GPT teacher head0.345
Teacher spread0.302 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainReproducibility
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of MicropalaeontologySame topicIsotope Analysis in EcologyFrench-language works237,207