MétaCan
Menu
Back to cohort
Record W2106994273 · doi:10.1093/ije/dyq111

DataSHIELD: resolving a conflict in contemporary bioscience--performing a pooled analysis of individual-level data without sharing the data

2010· article· en· W2106994273 on OpenAlexafffund
Michael Wolfson, Susan Wallace, Nicholas G. D. Masca, Geoffrey M. Rowe, Nuala A. Sheehan, Vincent Ferretti, Philippe Laflamme, Martin D. Tobin, John Macleod, Julian Little, Isabel Fortier, Bartha Maria Knoppers, Paul R. Burton

Bibliographic record

VenueInternational Journal of Epidemiology · 2010
Typearticle
Languageen
FieldEnvironmental Science
TopicHealth, Environment, Cognitive Aging
Canadian institutionsOntario Institute for Cancer ResearchStatistics CanadaThe Quebec Population Health Research NetworkMcGill Genome CentreUniversity of OttawaMcGill University and Génome Québec Innovation Centre
FundersMedical Research CouncilUniversity of LeicesterGenome CanadaNational Institute for Health and Care ResearchLeverhulme TrustBritish Heart FoundationWellcome Trust
KeywordsData scienceSample (material)Computer scienceFlexibility (engineering)Data sharingLegislationSet (abstract data type)Perspective (graphical)Sample size determinationManagement scienceData miningMedicineArtificial intelligenceEngineeringPolitical scienceStatisticsMathematicsLaw

Abstract

fetched live from OpenAlex

BACKGROUND: Contemporary bioscience sometimes demands vast sample sizes and there is often then no choice but to synthesize data across several studies and to undertake an appropriate pooled analysis. This same need is also faced in health-services and socio-economic research. When a pooled analysis is required, analytic efficiency and flexibility are often best served by combining the individual-level data from all sources and analysing them as a single large data set. But ethico-legal constraints, including the wording of consent forms and privacy legislation, often prohibit or discourage the sharing of individual-level data, particularly across national or other jurisdictional boundaries. This leads to a fundamental conflict in competing public goods: individual-level analysis is desirable from a scientific perspective, but is prevented by ethico-legal considerations that are entirely valid. METHODS: Data aggregation through anonymous summary-statistics from harmonized individual-level databases (DataSHIELD), provides a simple approach to analysing pooled data that circumvents this conflict. This is achieved via parallelized analysis and modern distributed computing and, in one key setting, takes advantage of the properties of the updating algorithm for generalized linear models (GLMs). RESULTS: The conceptual use of DataSHIELD is illustrated in two different settings. CONCLUSIONS: As the study of the aetiological architecture of chronic diseases advances to encompass more complex causal pathways-e.g. to include the joint effects of genes, lifestyle and environment-sample size requirements will increase further and the analysis of pooled individual-level data will become ever more important. An aim of this conceptual article is to encourage others to address the challenges and opportunities that DataSHIELD presents, and to explore potential extensions, for example to its use when different data sources hold different data on the same individuals.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.369
metaresearch head score (Gemma)0.658
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.993
Threshold uncertainty score0.778

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.3690.658
Meta-epidemiology (narrow)0.0020.004
Meta-epidemiology (broad)0.0040.003
Bibliometrics0.0140.029
Science and technology studies0.0040.020
Scholarly communication0.0200.023
Open science0.0070.025
Research integrity0.0070.012
Insufficient payload (model declined to judge)0.0160.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.334
GPT teacher head0.434
Teacher spread0.100 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainReproducibility
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations184
Published2010
Admission routes2
Has abstractyes

Explore more

Same venueInternational Journal of EpidemiologySame topicHealth, Environment, Cognitive AgingFrench-language works237,207