MétaCan
Menu
Back to cohort
Record W2944477024 · doi:10.23889/ijpds.v4i1.1103

Sharing linked data sets for research: results from a deliberative public engagement event in British Columbia, Canada

2019· article· en· W2944477024 on OpenAlexaffabout
Jack Teng, Colene Bentley, Michael Burgess, Kieran C. O’Doherty, Kimberlyn McGrail

Bibliographic record

VenueInternational Journal for Population Data Science · 2019
Typearticle
Languageen
FieldSocial Sciences
TopicData Analysis and Archiving
Canadian institutionsBC Centre for Disease ControlUniversity of GuelphCanadian Centre for Applied Research in Cancer ControlUniversity of British Columbia, Okanagan CampusUniversity of British Columbia
Fundersnot available
KeywordsDeliberationData sharingPublic relationsVotingEvent (particle physics)Public engagementInternet privacyData governanceProcess (computing)Political scienceData qualityComputer scienceBusinessMedicineLaw

Abstract

fetched live from OpenAlex

INTRODUCTION: Research using linked data sets can lead to new insights and discoveries that positively impact society. However, the use of linked data raises concerns relating to illegitimate use, privacy, and security (e.g., identity theft, marginalization of some groups). It is increasingly recognized that the public needs to be consulted to develop data access systems that consider both the potential benefits and risks of research. Indeed, there are examples of data sharing projects being derailed because of backlash in the absence of adequate consultation. (e.g., care.data in the UK). OBJECTIVES AND METHODS: This paper describes the results of a public deliberation event held in April 2018 in Vancouver, British Columbia. The purpose of this event was to develop informed and civic-minded public advice regarding the use and the sharing of linked data for research with a focus on the processes and regulations employed to release data. The event brought together 23 members of the public over two weekends. RESULTS: Participants developed and voted on 19 policy-relevant statements. Voting results and the rationale behind any disagreements are reported here. Taken together, these statements provide a broad view of public support and concerns regarding the use of linked data sets for research and offer guidance on measures that can be taken to improve the trustworthiness of policies and process around data sharing and use. CONCLUSIONS: Generally, participants were supportive of research using linked data because of the value they provide to society. Participants expressed a desire to see the data access request process made more efficient to facilitate more research, as long as there are adequate protections in place around security and privacy of the data.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.024
metaresearch head score (Gemma)0.041
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesOpen science
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: Qualitative
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.997
Threshold uncertainty score0.891

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0240.041
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.005
Science and technology studies0.0640.018
Scholarly communication0.0110.003
Open science0.0030.014
Research integrity0.0050.006
Insufficient payload (model declined to judge)0.0070.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.343
GPT teacher head0.486
Teacher spread0.143 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designQualitative
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations36
Published2019
Admission routes2
Has abstractyes

Explore more

Same venueInternational Journal for Population Data ScienceSame topicData Analysis and ArchivingFrench-language works237,207