Balancing Privacy and Utility in Secondary Data Use to Inform Policy
Bibliographic record
Abstract
ABSTRACT
 ObjectivesUsing existing data for research can generate new knowledge and evidence for policy with relatively little cost. Privacy concerns are paramount in such secondary usage of data collected on human subjects. Information privacy protection and data security are critical considerations in reuse and repurposing of data especially linked data, longitudinal data, and large amounts of data. Data sharing and privacy protection are both in the public interest and we need to assess the risk of “doing” (sharing) as well as the risk of “not doing” (not sharing or not protecting). 
 
 ApproachThe Alberta Centre for Child Family and Community Research (the Centre) establishes the Child Youth Data Lab that links and analyzes administrative data from multiple provincial ministries and the Child Data Centre of Alberta that repurposes research data and manages its access for reuse. The Centre partners with provincial Office of the Information Privacy Commissioner, Research Ethics Boards and leaders in the research communities and technology industry to design and develop measures to enable secondary use while safeguarding the data, and to explore and adopt best practices on data sharing processes, governance, and technologies.
 ResultsIn principle current privacy laws and regulations provide good guidance in collection, use, and disclosure of data, however there is a lack of consistency in the interpretation of these laws at the operational level with regard to secondary data use. The experiences of establishing different data sharing models at the Centre through multiple initiatives are discussed. Cross-sectoral broad partnership brings understanding and builds trusting relationships, which are crucial to establishing data sharing processes. The recognition of the significance of secondary data use to provide direction for policy and program development at the executive level provides commitment for data sharing initiatives. Strong governance structure consists multi-level ministry and multiple stakeholder involvement ensures ongoing support and engagement. The highest data security standards and anonymous solution for data linkage enables the sharing of data with good privacy protection.
 ConclusionSecondary use of data to improve system performance and contributing to scientific discovery has been broadly recognized. A balance between utility and privacy can be realized through broad partnership in building proper governance, technology, processes and policies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.123 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.006 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".