MétaCan
Menu
Back to cohort
Record W1917558546 · doi:10.1038/ejhg.2015.165

Harmonising and linking biomedical and clinical data across disparate data archives to enable integrative cross-biobank research

2015· article· en· W1917558546 on OpenAlexafffund
Ola Spjuth, Maria Krestyaninova, Janna Hastings, Huei-Yi Shen, Jani Heikkinen, Mélanie Waldenberger, Arnulf Langhammer, Claes Ladenvall, Tõnu Esko, Mats-Åke Persson, Jon Heggland, Joern Dietrich, Sandra Ose, Christian Gieger, Janina S. Ried, Annette Peters, Isabel Fortier, Eco J. C. de Geus, Jānis Kloviņš, Linda Zaharenko, Gonneke Willemsen, Jouke‐Jan Hottenga, Jan‐Eric Litton, Juha Karvanen, Dorret I. Boomsma, Leif Groop, Johan Rung, Juni Palmgren, Nancy L. Pedersen, Mark I. McCarthy, Cornelia M. van Duijn, Kristian Hveem, Andres Metspalu, Samuli Ripatti, Inga Prokopenko, Jennifer R. Harris

Bibliographic record

VenueEuropean Journal of Human Genetics · 2015
Typearticle
Languageen
FieldMedicine
TopicEthics in Clinical Research
Canadian institutionsMcGill University Health Centre
FundersMünchner Zentrum für GesundheitswissenschaftenTerveyden ja hyvinvoinnin laitosTartu ÜlikoolKarolinska InstitutetQueen's UniversityHelmholtz Zentrum MünchenNational Institute for Health and Care ResearchBundesministerium für Bildung und ForschungUniversity of OxfordQueen's University BelfastInnovative Medicines InitiativeImperial College LondonEuropean Federation of Pharmaceutical Industries and AssociationsLunds UniversitetKing's College LondonSwedish e-Science Research CentreOulun YliopistoEuropean CommissionBroad Institute
KeywordsBiobankData scienceDisparate systemSample (material)Computer scienceData accessSummitBiorepositoryData integrationData miningBioinformaticsDatabaseBiologyGeography

Abstract

fetched live from OpenAlex

A wealth of biospecimen samples are stored in modern globally distributed biobanks. Biomedical researchers worldwide need to be able to combine the available resources to improve the power of large-scale studies. A prerequisite for this effort is to be able to search and access phenotypic, clinical and other information about samples that are currently stored at biobanks in an integrated manner. However, privacy issues together with heterogeneous information systems and the lack of agreed-upon vocabularies have made specimen searching across multiple biobanks extremely challenging. We describe three case studies where we have linked samples and sample descriptions in order to facilitate global searching of available samples for research. The use cases include the ENGAGE (European Network for Genetic and Genomic Epidemiology) consortium comprising at least 39 cohorts, the SUMMIT (surrogate markers for micro- and macro-vascular hard endpoints for innovative diabetes tools) consortium and a pilot for data integration between a Swedish clinical health registry and a biobank. We used the Sample avAILability (SAIL) method for data linking: first, created harmonised variables and then annotated and made searchable information on the number of specimens available in individual biobanks for various phenotypic categories. By operating on this categorised availability data we sidestep many obstacles related to privacy that arise when handling real values and show that harmonised and annotated records about data availability across disparate biomedical archives provide a key methodological advance in pre-analysis exchange of information between biobanks, that is, during the project planning phase.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.049
metaresearch head score (Gemma)0.028
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Science and technology studies, Open science, Research integrity
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.414
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0490.028
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.003
Scholarly communication0.0010.000
Open science0.0020.008
Research integrity0.0000.005
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.878
GPT teacher head0.705
Teacher spread0.173 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations35
Published2015
Admission routes2
Has abstractyes

Explore more

Same venueEuropean Journal of Human GeneticsSame topicEthics in Clinical ResearchFrench-language works237,207