Developing a Health Data Analytic Service to Serve the Private Sector for Public Benefit
Bibliographic record
Abstract
ObjectivesOur Organization, a pan-Canadian network that facilitates publicly-funded multi-regional health research, is developing a process to undertake multi-jurisdictional data analytic services with private sector entities. We describe experiences and approaches to working with the private sector across Organization member sites, and concerns raised by our Public Advisory Committee. MethodsWe administered a survey with categorical and open-ended questions about work with the private sector including project volume, timelines, and differences from publicly-funded projects in terms of review and approval, fee structure, limitations on data access, and additional safeguards. It was distributed to 10 provincial and two pan-Canadian member sites holding linked health administrative data. Findings were organized using the Five Safes framework, a tool for evaluating access to privacy-sensitive data across five dimensions, and were presented at a quarterly PAC meeting to the group of ~15 members from diverse backgrounds. Key discussion and questions points were recorded. ResultsAmong 12 sites surveyed, 10 worked with the private sector in some capacity and 2 planned to in future. Common themes identified, though with variation across sites, included: safe data - some datasets are restricted from private sector use; safe projects - require public benefit and exclude market research, require research ethics board approval, monitor projects to ensure public benefit through milestone reporting; safe people and settings - no access to individual level data, require use of the site’s analytic services; safe outputs - require project summary/results be made public and shared with decision-makers, exclusion of site staff from authorship on publications. Other safeguards included dedicated access pathways to ensure publicly-funded work is not displaced. Most sites had higher project fees for the private sector. ConclusionGiven growing private sector demand for pan-Canadian, population-level, data analytic services to demonstrate product value and support adoption decisions, it is essential to set policies and practices that ensure public benefit. These results are supporting the launch of a pilot multi-site use case with a private sector organization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.094 | 0.106 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.013 | 0.004 |
| Scholarly communication | 0.010 | 0.007 |
| Open science | 0.006 | 0.015 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.020 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".