MétaCan
Menu
← Back to cohort
Record W214194437

Trends in collection, use and disclosure of personal information in contemporary health research: challenges for research governance.

2005· article· en· W214194437 on OpenAlexaffabout
Donald J. Willison

Bibliographic record

VenuePubMed · 2005
Typearticle
Languageen
FieldSocial Sciences
TopicHealth disparities and outcomes
Canadian institutionsSt. Joseph’s Healthcare Hamilton
Fundersnot available
KeywordsHealth careHealth policyHRHISPublic healthHealth equityEnvironmental healthPublic health informaticsHealth educationInternational healthPublic relationsObservational studyHealth services researchPsychologyMedicinePolitical scienceNursingEconomic growthEconomics
DOInot available

Abstract

fetched live from OpenAlex

Background: Changes in the Nature and Use of Personal Information Health Research Health research encompasses a heterogeneous set of research activities. This paper focuses on challenges that arise in the governance of observational research which is usually carried out without any direct contact between the researcher and the individuals being studied. Two broad areas of health research are heavily dependent on access to a wide range of existing person-level health information: 1. Public health, occupational health and safety, and the non-medical determinants of health and disease. The latter examines the relationship between health and lifestyle, environmental, and socioeconomic factors including income and education. Epidemiology is the foundation of much of this type of research. Research in this domain links both health and non-health information, such as occupation, education, and lifestyle information. 2. Health policy, health services research, and program evaluation examine the health care system and the effects of different policies and methods of health care delivery on the quality and efficiency of care provided. This type of research is informed by a wide variety of disciplines, including: economics, health policy, political sciences, sociology, anthropology, medicine, and epidemiology. Most health research requires person-level data, chiefly to increase precision in analysis. For example, when trying to determine the effect of exposure to an environmental toxin in a neighbourhood, with person-level data one can better examine the causal relationship by controlling for or holding constant known personal factors such as age and sex of the individual that relate to the outcome of interest. Similarly, when evaluating a policy to increase co-payments prescription drugs, it is prudent to examine across different income brackets the impact of that policy on the tendency to discontinue medications. In some cases, if using aggregate rather than individual-level data, it is possible to come to spurious conclusions about the effect of exposure (whether to a policy or an environmental toxin) on health outcomes. (1) Also, individual-level data are required to link information from disparate databases. This linkage creates the ability researchers to answer a much broader set of questions about the determinants of health, but it also raises major privacy concerns when these activities are being conducted without individual consent. Although data may be stripped of direct personal identifiers, the resultant records are often so rich in information that the residual risk of disclosure of identity through indirect means is sufficiently high that the data must be treated as if they were identifiable. In fact, with as little information as date-of-birth, sex, and full postal code, the majority of individuals in a particular region may be re-identified by linking with census tract information. (2) Trends in Data Collection, Use, Storage, and Disclosure Twenty years ago, only a handful of research centres across North America had the capacity to manipulate and link large data sets, and most government and other data repositories were used only claims adjudication. Medical records were all paper-based. Advances in the capacity of computers and the internet to store, manipulate and disseminate large amounts of data have changed dramatically the nature of collection, use and disclosure of personal information in contemporary health research. These advances have spawned two parallel developments: the planning and development of large disseminated health information networks that will serve multiple purposes beyond those direct clinical care; and the proliferation of decentralized holdings of personal data. In Canada, the United States, much of the European Union, Australia, and New Zealand, major efforts are underway to computerize patient records across health care settings, with the ability to share and link information from the records of physicians, diagnostic facilities, and health care institutions. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.535
metaresearch head score (Gemma)0.694
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.465
Threshold uncertainty score0.573

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5350.694
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0110.033
Science and technology studies0.0060.026
Scholarly communication0.0270.036
Open science0.0080.017
Research integrity0.0110.017
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.532
GPT teacher head0.467
Teacher spread0.064 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2005
Admission routes2
Has abstractyes

Explore more

Same venuePubMed→Same topicHealth disparities and outcomes→French-language works237,207→