MétaCan
Menu
Back to cohort
Record W4403846866 · doi:10.3897/biss.8.140268

What Matters for an occurrenceID and What Is an occurrenceID That Matters?

2024· article· en· W4403846866 on OpenAlexaff
Yi-Ming Gan, Abigail Benson, Emilio Mayorga, Jonathan Pye, Stephen K. Formel

Bibliographic record

VenueBiodiversity Information Science and Standards · 2024
Typearticle
Languageen
FieldComputer Science
TopicTime Series Analysis and Forecasting
Canadian institutionsOcean Tracking Network
Fundersnot available
KeywordsPolitical scienceEpistemologyPhilosophy

Abstract

fetched live from OpenAlex

In the Darwin Core data standard (Darwin Core Maintenance Group 2023), the concept of dwc:Occurrence (Wieczorek et al. 2012) in ecological data presents a data model construct not commonly found in data management practices used by ecological data collectors. We frequently encounter raw data without an Occurrence table. For example, the concept of Occurrence can be represented as a cell, where the rows represent sampling sites, the columns represent species, and the value of each cell indicates the count of the species at a specific site. A value in a cell in such a matrix can be interpreted as x number of individuals of species y occurred at sampling site z. While data providers tend to track data on tangible individual components (e.g., species, location, sample), generating "Occurrence records" typically requires pivoting and/or joining tables associated to these components. Maintaining a stable and persistent occurrenceID for an Occurrence record created through data transformation is not an easy task. This is especially true for long-term monitoring datasets, where the underlying tables used to generate Occurrence records are continuously updated. Additionally, most ecological data collectors are focused on the primary use of the data, not on the long term integration and accessibility of the data. The Occurrence concept is only required in data exchange format but not needed in ecological data management practices. The disconnect between the practical data management needs of data collectors and the abstractions required for data exchange raises challenges, particularly with an increasing expectation for globally unique and persistent occurrenceIDs. This presentation will explore the difficulties of creating and managing occurrenceIDs for data providers and managers, especially those who manage data using basic systems such as spreadsheets and simple relational databases. Maintaining stability and persistence of identifiers for inherently artificial constructs like Occurrences within the original, component-based data structure can pose significant challenges. We will explore why meaningful identifiers for occurrenceIDs are often preferred by data providers. We will unpack different use cases and delve into how and why occurrenceIDs were constructed for each use case. Through this discussion, we hope to spark a conversation that informs future data modeling efforts and addresses the inherent artificiality of Occurrences.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.050
metaresearch head score (Gemma)0.294
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.050
Threshold uncertainty score0.267

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0500.294
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0070.013
Science and technology studies0.0030.005
Scholarly communication0.0150.031
Open science0.0040.004
Research integrity0.0030.006
Insufficient payload (model declined to judge)0.0210.017

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.037
GPT teacher head0.283
Teacher spread0.246 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueBiodiversity Information Science and StandardsSame topicTime Series Analysis and ForecastingFrench-language works237,207