A review of a standards-based set of information security practices applied in Data Linkage Centers
Bibliographic record
Abstract
ABSTRACT ObjectiveThis research aimed to study regulatory and operational aspects related to information security, especially confidentiality, in organizations that systematically carry out record linkage. ApproachWe searched international experiences of data linkage units from the literature and from the catalog of International Population Data Linkage Network (IPDLN) members. In addition, we surveyed technical standards of the International Association for Standardization (ISO) on health informatics. ResultsWe studied organizations in Australia, Canada, UK and the United States. Six standards were selected for deep analysis. In the end, we organized a set of 75 practices relating to information security in data linkage units, grouped by 5 dimensions: infrastructure and operations; record linkage model; relationship with managers; relationship with researchers and relationship with the society. The linkage process must be described in a sufficiently clear and didactic way, so that ordinary citizens are able to understand that the privacy of their health information is protected. In addition to a transparent work process, the data linkage center must also make their privacy policies available. The Australian and Canadian experiences with ethic review committees that include social participation and awareness of media and explanations to the public are a good source of inspiration. Regarding safety, the institutions responsible for health databases should apply security controls in their information systems to consider the rules on consent to perform record linkage. Ideally, all institutions should seek full compliance with the controls recommended in the technical safety standards. However, the scarcity of resources (human, financial and technical) lead to the prioritization of the implementation of these security controls. The criteria for this prioritization can be given by feasibility analysis (cost / time impact, benefits), providing an orderly road map for the adoption of these measures. ConclusionThe practices systematized in this study can be used in order to check current information security conditions of data linkage centers and as guidelines for further improvements. This will certainly bring more confidence in the data linkage center process and, at least, help researchers, managers and society move forward toward the same objective of better public health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.078 | 0.165 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.034 | 0.040 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.009 | 0.009 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".