Public views and recommendations on the use of linked data for research: preliminary results from a public deliberation engagement
Bibliographic record
Abstract
IntroductionThe use of linked data for research is increasing, including in complexity of requests. Rules around access to and use of data necessarily trade-off risks related to privacy to achieve social benefits. Including informed and civic-minded public recommendations that consider different perspectives on privacy and benefit will improve related policy.
 Objectives and ApproachPopulation Data BC is conducting a deliberative public engagement regarding the use of complex linked data for research. Members of the public will be provided with written materials and hear speakers outlining considerations from multiple perspectives in data access and use, including benefits for health research, risks to privacy, and implications for disability and minority groups. Participants in the deliberation will then discuss questions about the use of linked data and ideas around principles for that use in small and large groups, and develop recommendations for data sharing policies.
 ResultsWe will be sharing our preliminary analysis of the public deliberation results at the conference. The public deliberation encourages the participants to develop policy recommendations that respect diversity of perspectives while negotiating constructive advice. It asks the group to make recommendations and to identify and explore issues on which the group has persistent disagreement. We will discuss insights into how the public values the use of data linkage and under what conditions such use becomes problematic. For example, we are hoping to gain insight about how publics determine if a project is in the public interest, or conversely, how a project may pose unacceptable harm.
 Conclusion/ImplicationsChanges in available data and increasing ability to link data makes it essential to include public views in systems of data access governance. Understanding the hopes and concerns of the public regarding the use of linked data for research will help develop data access regulations that reflect wide public interests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.047 | 0.107 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.010 |
| Open science | 0.007 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".