MétaCan
Menu
Retour à la cohorte
Enregistrement W2106994273 · doi:10.1093/ije/dyq111

DataSHIELD: resolving a conflict in contemporary bioscience--performing a pooled analysis of individual-level data without sharing the data

2010· article· en· W2106994273 sur OpenAlexafffund
Michael Wolfson, Susan Wallace, Nicholas G. D. Masca, Geoffrey M. Rowe, Nuala A. Sheehan, Vincent Ferretti, Philippe Laflamme, Martin D. Tobin, John Macleod, Julian Little, Isabel Fortier, Bartha Maria Knoppers, Paul R. Burton

Notice bibliographique

RevueInternational Journal of Epidemiology · 2010
Typearticle
Langueen
DomaineEnvironmental Science
ThématiqueHealth, Environment, Cognitive Aging
Établissements canadiensOntario Institute for Cancer ResearchStatistics CanadaThe Quebec Population Health Research NetworkMcGill Genome CentreUniversity of OttawaMcGill University and Génome Québec Innovation Centre
Organismes subventionnairesMedical Research CouncilUniversity of LeicesterGenome CanadaNational Institute for Health and Care ResearchLeverhulme TrustBritish Heart FoundationWellcome Trust
Mots-clésData scienceSample (material)Computer scienceFlexibility (engineering)Data sharingLegislationSet (abstract data type)Perspective (graphical)Sample size determinationManagement scienceData miningMedicineArtificial intelligenceEngineeringPolitical scienceStatisticsMathematicsLaw

Résumé

récupéré en direct d'OpenAlex

BACKGROUND: Contemporary bioscience sometimes demands vast sample sizes and there is often then no choice but to synthesize data across several studies and to undertake an appropriate pooled analysis. This same need is also faced in health-services and socio-economic research. When a pooled analysis is required, analytic efficiency and flexibility are often best served by combining the individual-level data from all sources and analysing them as a single large data set. But ethico-legal constraints, including the wording of consent forms and privacy legislation, often prohibit or discourage the sharing of individual-level data, particularly across national or other jurisdictional boundaries. This leads to a fundamental conflict in competing public goods: individual-level analysis is desirable from a scientific perspective, but is prevented by ethico-legal considerations that are entirely valid. METHODS: Data aggregation through anonymous summary-statistics from harmonized individual-level databases (DataSHIELD), provides a simple approach to analysing pooled data that circumvents this conflict. This is achieved via parallelized analysis and modern distributed computing and, in one key setting, takes advantage of the properties of the updating algorithm for generalized linear models (GLMs). RESULTS: The conceptual use of DataSHIELD is illustrated in two different settings. CONCLUSIONS: As the study of the aetiological architecture of chronic diseases advances to encompass more complex causal pathways-e.g. to include the joint effects of genes, lifestyle and environment-sample size requirements will increase further and the analysis of pooled individual-level data will become ever more important. An aim of this conceptual article is to encourage others to address the challenges and opportunities that DataSHIELD presents, and to explore potential extensions, for example to its use when different data sources hold different data on the same individuals.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,369
score de la tête « metaresearch » (Gemma)0,658
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Science ouverte
Catégories consensuellesMétarecherche
DomaineSignal candidat: Reproductibilité · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,993
Score d'incertitude au seuil0,778

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,3690,658
Méta-épidémiologie (sens strict)0,0020,004
Méta-épidémiologie (sens large)0,0040,003
Bibliométrie0,0140,029
Études des sciences et des technologies0,0040,020
Communication savante0,0200,023
Science ouverte0,0070,025
Intégrité de la recherche0,0070,012
Charge utile insuffisante (le modèle a refusé de juger)0,0160,008

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,334
Tête enseignante GPT0,434
Écart entre enseignants0,100 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.

Devis d'étudeThéorique ou conceptuel
DomaineReproductibilité
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations184
Publié2010
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueInternational Journal of EpidemiologyMême sujetHealth, Environment, Cognitive AgingTravaux en français237 207