MétaCan
Menu
Back to cohort
Record W4416854882 · doi:10.3897/biss.9.180398

Tracking Natural Science Objects and Their Physical and Digital Derivatives in DINA

2025· article· en· W4416854882 on OpenAlexaff
Jonas Grieb, James Macklin, David Peter Shorthouse, Christian Bölling, Volker Lohrmann, Michaela Grein, Etta Grotrian, Satpal Bilkhu, Claus Weiland

Bibliographic record

VenueBiodiversity Information Science and Standards · 2025
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsAgriculture and Agri-Food Canada
Fundersnot available
KeywordsMetadataContext (archaeology)IdentifierTimestampTracking (education)Tracking systemDigital preservationData collectionUnique identifierNatural (archaeology)

Abstract

fetched live from OpenAlex

Provenance plays an important role in natural history collections, but capturing this information accurately, proved challenging for legacy digital collection management systems. DINA*1 is being developed to address these limitations through a process-oriented data model that more effectively captures the complexity and context of provenance information (Bölling et al. 2022). DINA is an open-source, robust sample- and specimen-based collection management system in production and developed by an unincorporated international consortium of technologists and practitioners of the natural sciences (Glöckler et al. 2020). Its governance and data models help to foster the adoption of FAIR principles (Findable, Accessible, Interopable, Reusable) and to integrate objects across science domains. DINA's innovative "samplistic" data model records metadata on stepwise, hierarchical processes that generate physical and digital derivatives from parent material samples, accommodating complex real-life sample trajectories, for example: A fossil specimen is acquired by a museum, which contains the remains of several organisms in hardened resin. Later, it is sawed into smaller pieces; some pieces are stored under new catalog numbers, a tiny piece is sent for a destructive C14 (radioactive carbon) analysis (Fig. 1 a). A naturally deceased individual of a mammal species is collected under a material transfer agreement (MTA). Later, subsamples like teeth, bones, and tissues undergo different preparation and preservation processes and are finally stored in specialized collections, each with its own identifier (scheme); a subsample might even be retrieved for sequencing. Strong provenance tracking is required to ensure that regulatory constraints like the MTA are consistently passed down to any derivatives. A soil core is collected from a sampling site and registered together with observational metadata according to the MIxS (Minimum Information about Any Sequence) soil extension standard. Several subsamples are extracted before the remaining core is preserved. The subsamples are subjected to DNA extraction and sequencing to analyze environmental DNA (eDNA). The resulting sequences are then processed computationally to identify the species present (Fig. 1 b). A fossil specimen is acquired by a museum, which contains the remains of several organisms in hardened resin. Later, it is sawed into smaller pieces; some pieces are stored under new catalog numbers, a tiny piece is sent for a destructive C14 (radioactive carbon) analysis (Fig. 1 a). A naturally deceased individual of a mammal species is collected under a material transfer agreement (MTA). Later, subsamples like teeth, bones, and tissues undergo different preparation and preservation processes and are finally stored in specialized collections, each with its own identifier (scheme); a subsample might even be retrieved for sequencing. Strong provenance tracking is required to ensure that regulatory constraints like the MTA are consistently passed down to any derivatives. A soil core is collected from a sampling site and registered together with observational metadata according to the MIxS (Minimum Information about Any Sequence) soil extension standard. Several subsamples are extracted before the remaining core is preserved. The subsamples are subjected to DNA extraction and sequencing to analyze environmental DNA (eDNA). The resulting sequences are then processed computationally to identify the species present (Fig. 1 b). In all examples, the processing history of each subsample can be modeled in DINA so provenance can always be traced back to the original sample. Knowledge from any derivative analysis can be linked back to the parent sample and other derivatives. This ensures that regulatory restrictions applying to a parent material sample are passed down consistently. Besides parent-child relationships, the model can represent other associations, e.g., host, parasite, and vector relationships linking samples in different collections. Each sample in the provenance chain, as well as first-class related objects like projects, collections, events, people, protocols, and storage, has a system-defined globally unique identifier. DINA integrates with external systems where possible via persistent identifiers, including reuse of scientific names from community-curated biodiversity sources through the Global Names Architecture. DINA’s application programming interface simplifies data import and migration and can be used by scientific programming languages like Python or R to access data objects or their relationships. Flexible user-defined “managed attributes” enhance object and derivative metadata when standards do not yet exist. When standards like MIxS or MIDS (Minimum Information about a Digital Specimen) exist, these are incorporated as field extensions. Looking ahead, we aim to further strengthen linkages and provenance tracking in DINA, extending support to model complex bio-geo relationships and enhancing linkage of material samples to external resources such as publications.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication
Consensus categoriesScholarly communication
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.749
Threshold uncertainty score0.994

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.002
Science and technology studies0.0010.002
Scholarly communication0.0070.126
Open science0.0010.002
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.027
GPT teacher head0.316
Teacher spread0.289 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBiodiversity Information Science and StandardsSame topicResearch Data Management PracticesFrench-language works237,207