Political and Policy Arguments for Integrated Data
Bibliographic record
Abstract
ABSTRACT IntroductionThere is little argument that integrated data can provide a valuable resource for improved health system management, planning, and accountability as well as discovery and commercial use, but policies to enable and support integrated data fall short of the potential represented by integrated data. To understand the current level of progress on policy for integrated data, we looked at two successful and two unsuccessful efforts to support the creation and use of integrated data in health systems. Methods/ApproachWe used document and literature analysis to develop descriptions of the Icelandic Health Sector Database Act, the creation of the Institute for Clinical Evaluative Sciences in Ontario (Canada), the care.data initiative in the United Kingdom, and the Health Datapalooza initiative in the US and used an Ideas, Institutions and Actors framework to compare the experience with integrated data policy and politics. Results and discussionOur analysis suggests that institutions around integrated data remain under-developed and largely focused on specific aspects of integrated data policy or use. There are at least two sets of dominant ideas around integrated data – data as a tool for economic development and health system performance and data as a threat to privacy and liberty – that are often diametrically opposed in different jurisdictions. To a great extent, powerful actors remain disengaged from integrated data discussions and leadership engaged in integrated data policy and politics remains isolated from larger policy and political discussions. The medical profession along with civil society groups can mount effective opposition to integrated data initiatives, although potentially for different reasons (accountability and privacy concerns respectively). ConclusionsOur analysis suggests several key issues around successful integrated data policy and politics that support the importance of strong leadership, an incremental approach to institution building that focuses on public benefits, strongly alignment to missions that are congruent with societal values, and stronger attention to effective and rapid implementation of policy. In addition to the cases studied here, the success of smaller sub-national (e.g. state or provincial) efforts suggests that smaller efforts tend to work better although their success may not receive the attention that could support larger efforts to integrate data on the national level. Further work should focus chiefly on the extension of these arguments to non-health sectors to realize the full value of integrated data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.205 | 0.184 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.016 | 0.073 |
| Scholarly communication | 0.050 | 0.042 |
| Open science | 0.005 | 0.023 |
| Research integrity | 0.027 | 0.032 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".