MétaCan
Menu
Back to cohort
Record W7061054402

Overcoming data privacy and data gravity challenges in bioinformatics research

2023· other· en· W7061054402 on OpenAlexaboutno aff

Bibliographic record

VenueeScholarship (California Digital Library) · 2023
Typeother
Languageen
FieldEngineering
TopicParticle accelerators and beam dynamics
Canadian institutionsnot available
FundersNational Human Genome Research InstituteNational Heart, Lung, and Blood InstituteScience Mission DirectorateNational Cancer InstituteNational Institutes of HealthJapan Agency for Medical Research and DevelopmentNational Aeronautics and Space Administration
KeywordsBiobankInformation privacyData sharingData anonymizationPrecision medicineBig dataData integrationMedical researchData modeling
DOInot available

Abstract

fetched live from OpenAlex

Next-generation sequencing technologies have generated a massive amount of DNA, RNA, and protein sequences since their inception. However, data privacy policies often restrict sharing such data for the risk of re-identifying individuals from whom the sequences were generated. Even when all the data from a sequencing experiment is available, it is often insufficient for statistical power or training machine learning models. Despite the lack of data, sometimes the data sets are ironically too large to realistically share with researchers. In this thesis, I explore methods to overcome challenges of data privacy and data gravity in bioinformatics research. In collaboration with QIMR Berghofer and the Riken Center for Integrative Medical Sciences, we used federated methods to analyze genomic data from the BioBank Japan in situ to classify variants of uncertain significance while preserving privacy. With the Department of Laboratory Medicine and Pathology at the University of Washington, we developed a statistical model that demonstrates how using responsibly shared clinical evidence alone can classify variants of uncertain significance which occur at the rate of 1 in 100,000 people within just a few years. With researchers from McGill University, we reviewed the state of the art in federated computing technologies and how well they satisfy the privacy restrictions from the General Data Protection Regulation. With researchers from NASA, Amazon, and Intel, we developed a federated learning framework to run between terrestrial and space-borne compute infrastructure, laying the groundwork for subsequent experiments which preclude the need to transfer large datasets across astronomical distances. Finally, at NASA, we used a causal inference machine learning ensemble to infer robust correlation between mouse liver gene expression and a corresponding lipid density phenotype in space-flown mice.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.150
metaresearch head score (Gemma)0.247
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.150
Threshold uncertainty score0.792

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1500.247
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0040.009
Science and technology studies0.0050.020
Scholarly communication0.0150.039
Open science0.0060.020
Research integrity0.0070.016
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.127
GPT teacher head0.310
Teacher spread0.184 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueeScholarship (California Digital Library)Same topicParticle accelerators and beam dynamicsFrench-language works237,207