MétaCan
Menu
Back to cohort
Record W2128349406 · doi:10.1093/ije/dyu188

DataSHIELD: taking the analysis to the data, not the data to the analysis

2014· article· en· W2128349406 on OpenAlexaff
Amadou Gaye, Yannick Marcon, Julia Isaeva, Philippe Laflamme, Andrew Turner, Elinor M Jones, Joel T. Minion, Andy Boyd, Chris Newby, Marja-Liisa Nuotio, Rebecca Wilson, O. W. Butters, Barnaby Murtagh, İpek Demir, Dany Doiron, Lisette Giepmans, Susan Wallace, Isabelle Budin‐Ljøsne, Carsten Oliver Schmidt, Paolo Boffetta, Mathieu Boniol, Maria Bota, Kim W. Carter, Chris Dibben, Richard W. Francis, Tero Hiekkalinna, Kristian Hveem, Kirsti Kvaløy, Seán Millar, Ivan J. Perry, Annette Peters, Catherine M. Phillips, Frank Popham, Gillian Raab, Eva Reischl, Nuala A. Sheehan, Mélanie Waldenberger, Markus Perola, Edwin R. van den Heuvel, John Macleod, Bartha Maria Knoppers, Ronald P. Stolk, Isabel Fortier, Jennifer R. Harris, Bruce HR Woffenbuttel, Madeleine J. Murtagh, Vincent Ferretti, Paul R. Burton

Bibliographic record

VenueInternational Journal of Epidemiology · 2014
Typearticle
Languageen
FieldEnvironmental Science
TopicHealth, Environment, Cognitive Aging
Canadian institutionsResearch CanadaOntario Institute for Cancer ResearchOntario GenomicsMcGill University Health Centre
FundersEconomic and Social Research CouncilMedical Research CouncilEuropean CommissionWellcome Trust
KeywordsPoolingComputer scienceData scienceConfidentialityBiomedicineIntellectual propertySample (material)Health careComputer securityPolitical science

Abstract

fetched live from OpenAlex

BACKGROUND: Research in modern biomedicine and social science requires sample sizes so large that they can often only be achieved through a pooled co-analysis of data from several studies. But the pooling of information from individuals in a central database that may be queried by researchers raises important ethico-legal questions and can be controversial. In the UK this has been highlighted by recent debate and controversy relating to the UK's proposed 'care.data' initiative, and these issues reflect important societal and professional concerns about privacy, confidentiality and intellectual property. DataSHIELD provides a novel technological solution that can circumvent some of the most basic challenges in facilitating the access of researchers and other healthcare professionals to individual-level data. METHODS: Commands are sent from a central analysis computer (AC) to several data computers (DCs) storing the data to be co-analysed. The data sets are analysed simultaneously but in parallel. The separate parallelized analyses are linked by non-disclosive summary statistics and commands transmitted back and forth between the DCs and the AC. This paper describes the technical implementation of DataSHIELD using a modified R statistical environment linked to an Opal database deployed behind the computer firewall of each DC. Analysis is controlled through a standard R environment at the AC. RESULTS: Based on this Opal/R implementation, DataSHIELD is currently used by the Healthy Obese Project and the Environmental Core Project (BioSHaRE-EU) for the federated analysis of 10 data sets across eight European countries, and this illustrates the opportunities and challenges presented by the DataSHIELD approach. CONCLUSIONS: DataSHIELD facilitates important research in settings where: (i) a co-analysis of individual-level data from several studies is scientifically necessary but governance restrictions prohibit the release or sharing of some of the required data, and/or render data access unacceptably slow; (ii) a research group (e.g. in a developing nation) is particularly vulnerable to loss of intellectual property-the researchers want to fully share the information held in their data with national and international collaborators, but do not wish to hand over the physical data themselves; and (iii) a data set is to be included in an individual-level co-analysis but the physical size of the data precludes direct transfer to a new site for analysis.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.105
metaresearch head score (Gemma)0.372
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.994
Threshold uncertainty score0.555

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1050.372
Meta-epidemiology (narrow)0.0030.003
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0080.012
Science and technology studies0.0020.010
Scholarly communication0.0130.013
Open science0.0060.012
Research integrity0.0030.009
Insufficient payload (model declined to judge)0.0500.033

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.186
GPT teacher head0.428
Teacher spread0.242 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainReproducibility
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations306
Published2014
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of EpidemiologySame topicHealth, Environment, Cognitive AgingFrench-language works237,207