MétaCan
Menu
Retour à la cohorte
Enregistrement W4416854882 · doi:10.3897/biss.9.180398

Tracking Natural Science Objects and Their Physical and Digital Derivatives in DINA

2025· article· en· W4416854882 sur OpenAlexaff
Jonas Grieb, James Macklin, David Peter Shorthouse, Christian Bölling, Volker Lohrmann, Michaela Grein, Etta Grotrian, Satpal Bilkhu, Claus Weiland

Notice bibliographique

RevueBiodiversity Information Science and Standards · 2025
Typearticle
Langueen
DomaineComputer Science
ThématiqueResearch Data Management Practices
Établissements canadiensAgriculture and Agri-Food Canada
Organismes subventionnairesnon disponible
Mots-clésMetadataContext (archaeology)IdentifierTimestampTracking (education)Tracking systemDigital preservationData collectionUnique identifierNatural (archaeology)

Résumé

récupéré en direct d'OpenAlex

Provenance plays an important role in natural history collections, but capturing this information accurately, proved challenging for legacy digital collection management systems. DINA*1 is being developed to address these limitations through a process-oriented data model that more effectively captures the complexity and context of provenance information (Bölling et al. 2022). DINA is an open-source, robust sample- and specimen-based collection management system in production and developed by an unincorporated international consortium of technologists and practitioners of the natural sciences (Glöckler et al. 2020). Its governance and data models help to foster the adoption of FAIR principles (Findable, Accessible, Interopable, Reusable) and to integrate objects across science domains. DINA's innovative "samplistic" data model records metadata on stepwise, hierarchical processes that generate physical and digital derivatives from parent material samples, accommodating complex real-life sample trajectories, for example: A fossil specimen is acquired by a museum, which contains the remains of several organisms in hardened resin. Later, it is sawed into smaller pieces; some pieces are stored under new catalog numbers, a tiny piece is sent for a destructive C14 (radioactive carbon) analysis (Fig. 1 a). A naturally deceased individual of a mammal species is collected under a material transfer agreement (MTA). Later, subsamples like teeth, bones, and tissues undergo different preparation and preservation processes and are finally stored in specialized collections, each with its own identifier (scheme); a subsample might even be retrieved for sequencing. Strong provenance tracking is required to ensure that regulatory constraints like the MTA are consistently passed down to any derivatives. A soil core is collected from a sampling site and registered together with observational metadata according to the MIxS (Minimum Information about Any Sequence) soil extension standard. Several subsamples are extracted before the remaining core is preserved. The subsamples are subjected to DNA extraction and sequencing to analyze environmental DNA (eDNA). The resulting sequences are then processed computationally to identify the species present (Fig. 1 b). A fossil specimen is acquired by a museum, which contains the remains of several organisms in hardened resin. Later, it is sawed into smaller pieces; some pieces are stored under new catalog numbers, a tiny piece is sent for a destructive C14 (radioactive carbon) analysis (Fig. 1 a). A naturally deceased individual of a mammal species is collected under a material transfer agreement (MTA). Later, subsamples like teeth, bones, and tissues undergo different preparation and preservation processes and are finally stored in specialized collections, each with its own identifier (scheme); a subsample might even be retrieved for sequencing. Strong provenance tracking is required to ensure that regulatory constraints like the MTA are consistently passed down to any derivatives. A soil core is collected from a sampling site and registered together with observational metadata according to the MIxS (Minimum Information about Any Sequence) soil extension standard. Several subsamples are extracted before the remaining core is preserved. The subsamples are subjected to DNA extraction and sequencing to analyze environmental DNA (eDNA). The resulting sequences are then processed computationally to identify the species present (Fig. 1 b). In all examples, the processing history of each subsample can be modeled in DINA so provenance can always be traced back to the original sample. Knowledge from any derivative analysis can be linked back to the parent sample and other derivatives. This ensures that regulatory restrictions applying to a parent material sample are passed down consistently. Besides parent-child relationships, the model can represent other associations, e.g., host, parasite, and vector relationships linking samples in different collections. Each sample in the provenance chain, as well as first-class related objects like projects, collections, events, people, protocols, and storage, has a system-defined globally unique identifier. DINA integrates with external systems where possible via persistent identifiers, including reuse of scientific names from community-curated biodiversity sources through the Global Names Architecture. DINA’s application programming interface simplifies data import and migration and can be used by scientific programming languages like Python or R to access data objects or their relationships. Flexible user-defined “managed attributes” enhance object and derivative metadata when standards do not yet exist. When standards like MIxS or MIDS (Minimum Information about a Digital Specimen) exist, these are incorporated as field extensions. Looking ahead, we aim to further strengthen linkages and provenance tracking in DINA, extending support to model complex bio-geo relationships and enhancing linkage of material samples to external resources such as publications.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,002
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante
Catégories consensuellesCommunication savante
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,749
Score d'incertitude au seuil0,994

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0020,002
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0010,002
Études des sciences et des technologies0,0010,002
Communication savante0,0070,126
Science ouverte0,0010,002
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,027
Tête enseignante GPT0,316
Écart entre enseignants0,289 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBiodiversity Information Science and StandardsMême sujetResearch Data Management PracticesTravaux en français237 207