MétaCan
Menu
Back to cohort
Record W4410836337 · doi:10.5194/jm-44-145-2025

Community guidelines to increase the reusability of marine microfossil assemblage data

2025· article· en· W4410836337 on OpenAlexaff
Lukas Jonkers, Anne Strack, Montserrat Alonso‐García, Simon D’haenens, Robert Huber, Michal Kučera, Iván Hernández‐Almeida, Chloe Jones, Brett Metcalfe, Rajeev Saraswat, Lóránd Silye, Sanjay Kumar Verma, Muhamad Naim Abd Malek, Gerald Auer, Catia F Barbosa, María Ángeles Bárcena, Karl‐Heinz Baumann, Flavia Boscolo‐Galazzo, Joeven Austine S. Calvelo, Lucilla Capotondi, Martina Caratelli, Jorge Cardich, Humberto Carvajal-Chitty, Markéta Chroustová, Helen K. Coxall, Renata M. de Mello, Anne de Vernal, Paula Diz, Kirsty M. Edgar, Helena L. Filipsson, Ángela Fraguas, Heather Furlong, Giacomo Galli, Natalia García Chapori, Robyn Granger, Jeroen Groeneveld, Adil Imam, Rebecca Jackson, David Lazarus, Julie Meilland, Marína Molčan Matejová, Raphaël Morard, Caterina Morigi, Sven N. Nielsen, Diana Ochoa, Maria Rose Petrizzo, Andrés S. Rigual‐Hernández, Marina C. Rillo, Matthew L. Staitis, Gamze Tanık, Raúl Tapia, Nishant Vats, Bridget S. Wade, Alexander Weinmann

Bibliographic record

VenueJournal of Micropalaeontology · 2025
Typearticle
Languageen
FieldEnvironmental Science
TopicIsotope Analysis in Ecology
Canadian institutionsUniversité du Québec à Montréal
FundersNederlandse Organisatie voor Wetenschappelijk OnderzoekDeutsche Forschungsgemeinschaft
KeywordsReusabilityAssemblage (archaeology)OceanographyGeologyEnvironmental resource managementEnvironmental scienceComputer sciencePaleontology

Abstract

fetched live from OpenAlex

Abstract. Data on marine microfossil assemblage composition have multiple applications. Initially, they were primarily used for (chrono)stratigraphy and palaeoecology, but these data are now also widely used to study evolutionary and ecological processes, such as past biodiversity and its links with environmental dynamics, or to provide a basis for conservation efforts and biomonitoring. The large range of potential applications renders microfossil abundance data ideal for reuse. However, the complexity inherent in taxonomic data, which encompass extant and extinct species, coupled with the inherent intricacies of information on biological communities extracted from sedimentary archives, poses considerable hurdles in reusing marine microfossil data, even when they are publicly available. Here, we present guidelines derived from an online survey conducted within the marine micropalaeontological community, aimed at improving the reusability of microfossil assemblage data. These guidelines advocate for clarity and transparency in the documentation of the methods and the outcome, and we outline the data attributes required for effective reuse of micropalaeontological data. These guidelines are intended for researchers who generate microfossil abundance datasets and for reviewers, editors, and data curators at repositories. A total of 113 researchers evaluated the relevance of about 50 data attributes that might be needed to enable and maximise the reuse of marine microfossil abundance datasets. Each property is ranked based on the survey results. All information is, in principle, considered “desired”. Information that improves the reusability is ranked as “recommended”, and information that is required for reuse is ranked as “essential”. Analysis of a selection of datasets available online reveals a rather large gap between data properties deemed essential by survey participants and what is actually contained in publicly available microfossil assemblage datasets. While the survey indicates that the micropalaeontological community values good data stewardship, improving data reusability still requires new efforts to incorporate all the essential information. The guidelines presented here are intended as a step in that direction. Determining the optimal forms and formats for data sharing are obvious next steps the community needs to take.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.004
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.175
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0050.004
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0030.003
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.043
GPT teacher head0.345
Teacher spread0.302 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of MicropalaeontologySame topicIsotope Analysis in EcologyFrench-language works237,207