Learning from experience: privacy and the secondary use of data in health research
Bibliographic record
Abstract
Health services research must continually address the question: Under what conditions may data not collected specifically for research, such as primary medical data, be re-used for research without compromising the privacy of the data-subjects? For secondary use of data in research there are basically three options. Option A: Use personal data with consent or other assent from the data-subjects. To make this both fairer and more practical, in many circumstances broader construals of consent, or permission or approval, need to be explored and instituted. Option B: Anonymise the data, then use them. For many studies, this is the most practical and desirable option. The craft of anonymisation, including reversible anonymisation, or key-coding, needs to be developed and more fully supported under law. Option C: Use personal data without explicit consent, under a public interest mandate. Whether and how the data should be anonymised will depend on the situation. Public health mandates and protections deserve to be clarified, strengthened and extended for a variety of surveillance, registration, clinical audit, health services research and other types of investigation. Safeguards are an integral part of the research promise to the public, offer crucial reassurance and should be emphasised. For health services research, databases are core resources, and their stewardship must be cultivated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.358 | 0.394 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.014 | 0.184 |
| Scholarly communication | 0.034 | 0.081 |
| Open science | 0.005 | 0.035 |
| Research integrity | 0.018 | 0.025 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".