MétaCan
Menu
Back to cohort
Record W2606810975 · doi:10.23889/ijpds.v1i1.364

Big Data, Big Responsibility! Building best-practice privacy strategies into a large-scale neuroinformatics platform

2017· article· en· W2606810975 on OpenAlexaffabout
Christina Popovich, Francis Jeanson, Brendan Behan, Shannon Lefaivre, Aparna Wagle Shukla

Bibliographic record

VenueInternational Journal for Population Data Science · 2017
Typearticle
Languageen
FieldEnvironmental Science
TopicHealth, Environment, Cognitive Aging
Canadian institutionsOntario Brain Institute
Fundersnot available
KeywordsNeuroinformaticsComputer scienceData sharingBig dataData scienceHealth informaticsData governanceInformaticsVariety (cybernetics)Best practiceAnalyticsComputer securityHealth careMedicineEngineeringData miningPolitical scienceArtificial intelligence

Abstract

fetched live from OpenAlex

ABSTRACT ObjectiveThe Ontario Brain Institute (OBI) has begun to catalyze scientific discovery in the field of neuroscience through its’ large-scale informatics platform, known as Brain-CODE (Centre for Ontario Data Exploration). Brain-CODE manages the acquisition, storage, processing, and analytics of multidimensional data collected from patients with a variety of brain disorders. Our vision is for the platform to act as an informatics catalyst; encouraging multidisciplinary research collaboration, data integration, and innovation in neuroscience research. Brain-CODE’s infrastructure was designed with best-practice privacy strategies built at the forefront to enable secure data capture of sensitive patient information in a manner that abides by government legislation while fostering data sharing and linking opportunities. ApproachPrivacy and security features have been incorporated into the very foundation of Brain-CODE’s comprehensive guidelines, which are reinforced by our state-of-the-art approaches to keep patient data safe. To ensure clarity for study participants, we have developed standard consent language outlining how sensitive patient data will be collected, entered, de-identified, and shared using Brain-CODE. Moreover, our tiered approach to data accessibility enables the storage of encrypted Ontario Health Card Numbers as well as other patient information, secure long-term storage of de-identified data, and data sharing opportunities by request from third parties following risk-based analysis re-identification techniques. OBI has also established a comprehensive Information Security Policy and Informatics Governance Policies, as well as a carried out a Privacy Impact Assessment and Threat Risk Assessment for Brain-CODE. ResultsBrain-CODE is proudly named a "Privacy by Design" Ambassador by the Office of the Information and Privacy Commissioner of Ontario, Canada. Moreover, approximately 200 neuroscience researchers and 35 institutions from across Canada have adopted our standard consent language to enable secure data sharing within and across neurological disorders as well as linkage opportunities with national and international databases in a secure environment. ConclusionOBI’s rigorous approach to data sharing in the field of neuroscience maintains the accessibility of research data for big discoveries without compromising patient privacy and security. We believe that Brain-CODE is a powerful and advantageous tool; moving neuroscience research from independent silos to an integrative system approach for improving patient health. OBI’s vision for improved brain health for patients living with neurological disorders paired with Brain-CODE’s best-practice strategies in privacy protection of patient data offer a novel and innovative approach to “big data” initiatives aimed towards improving public health and society world-wide.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.101
metaresearch head score (Gemma)0.114
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.899
Threshold uncertainty score0.535

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1010.114
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0080.014
Scholarly communication0.0240.035
Open science0.0070.037
Research integrity0.0070.011
Insufficient payload (model declined to judge)0.0080.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.136
GPT teacher head0.427
Teacher spread0.292 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2017
Admission routes2
Has abstractyes

Explore more

Same venueInternational Journal for Population Data ScienceSame topicHealth, Environment, Cognitive AgingFrench-language works237,207