MétaCan
Menu
Back to cohort
Record W6931687237 · doi:10.5281/zenodo.6921457

Diary of our initiatory journey on the continent of data citation in SSH

2022· article· en· W6931687237 on OpenAlexaff

Bibliographic record

VenueISTI Open Portal · 2022
Typearticle
Languageen
FieldEnvironmental Science
TopicMarine and fisheries research
Canadian institutionsHumber Polytechnic
FundersHorizon 2020 Framework Programme
KeywordsCitationWork (physics)Table (database)Set (abstract data type)Quality (philosophy)Metaphor

Abstract

fetched live from OpenAlex

If citation is a common practice for publications, it is relatively new for data especially in SSH. This paper will present the work carried out during the SSHOC project about data citation in general and more precisely how to make them actionable. The metaphor of a travel journal of an expedition seemed appropriate to us to present this work carried out during the SSHOC project. The first part was to study this terra incognita by making an inventory of citation practices (https://doi.org/10.5281/zenodo.3595965). To summarize, we discovered that in the research communities we investigated, practices were seldom standardized and were very diverse, generally producing citations that could not be processed by machines: in other words they were not “actionable”. This led us to develop a sort of guide necessary to journey through this new, uncharted territory in the form of a set of recommendations ( https://doi.org/10.5281/zenodo.5361717) to build citations in SSH. So as not to reinvent the wheel, we based these recommendations on existing principles created by Force11 ( https://doi.org/10.25490/a97f-egyk) by adapting them to the specific characteristics of the SSH data. These recommendations were validated by a committee of experts from different backgrounds and structures (RDA participants, CODATA director, OpenAire Engineers etc.) during a round table (https://www.sshopencloud.eu/news/roundtable-experts-data-citation) and in a parallel review process. Then we decided to analyze the resources available in this new territory, that is, the repositories that are so crucial to be able to cite data. We carried out an analysis of 85 repositories against 7 quality criteria based on the recommendations which ensure continuity with the work mentioned above: PID from “Unique Identification & Persistence” Landing page from “Access” Structured metadata from “Importance & Credit and Attribution” Cite as from “Evidence, Specificity & Verifiability” Versioning from “Specificity and Verifiability” Standardized vocabularies from “Interoperability and Flexibility” Links to publications from “Importance” The results of this survey (https://doi.org/10.5281/zenodo.5603306) are encouraging - even if there is room for improvement, particularly in the use of Persistent Identifiers. Importantly, the presence of a landing page in almost all cases allowed us to build up a test sample made up of a very diverse dataset from those repositories for which we want to build standardized and actionable citations. In parallel we developed a tool in order to “harvest” the resources found in this new land so as to better understand them and also be able to explain them to others. We developed a prototype composed of three components: a harvester which grabs information about a dataset and normalizes it an API to disseminate the metadata of the citation thereby making it actionable a citation viewer for human purposes For the first iteration to populate this prototype, we used the dataset collected during our survey of repositories and we are going to gradually add more datasets from various sources. This prototype is primarily designed to implement what we called “actionability” to a citation and provide a ready-to-use citation in various citation formats. Starting from the PID of a dataset, the prototype attempts to aggregate metadata from different sources: the repository of the dataset, the PID Registration Agency and a number of Knowledge Graphs. For instance, while metadata associated with a DOI (Digital Object Identifier) are limited and those provided by a handle are even more scarce, it is possible to get more information from a landing page and thus enrich the citation. We also used another indirect approach to gather additional information by using a registry of repositories (RE3Data https://www.re3data.org/) which provides, among other things, information on the available APIs available for a specific repository. Thus the prototype can give a unified view of information about datasets coming from different sources. For researchers, it thus avoids cumbersome work on how to cite a dataset or get information about its provenance. In return, it makes a researcher aware of the importance of properly documenting a dataset and depositing it in a “good” repository. This paper will present in greater detail what we learned at each step of this expedition and how a research project can take advantage of a good citation system to enhance the visibility of the output. We will also introduce the potential uses based on the information provided by the prototype such as the possibility of associating a specific tool to process data or the use of this information as a base to build data papers.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.102
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Scholarly communication
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: Qualitative
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.980
Threshold uncertainty score0.116

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.102
Meta-epidemiology (narrow)0.0000.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0070.013
Science and technology studies0.0190.010
Scholarly communication0.0200.020
Open science0.0020.013
Research integrity0.0040.012
Insufficient payload (model declined to judge)0.0160.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.131
GPT teacher head0.349
Teacher spread0.219 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designQualitative
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueISTI Open PortalSame topicMarine and fisheries researchFrench-language works237,207