MétaCan
Menu
Back to cohort
Record W4367337274 · doi:10.2196/45823

Comparing Decentralized Learning Methods for Health Data Models to Nondecentralized Alternatives: Protocol for a Systematic Review

2023· review· en· W4367337274 on OpenAlexvenueno aff
José Miguel Diniz, Henrique Vasconcelos, Júlio Souza, Rita Rb-Silva, Carolina Ameijeiras-Rodriguez, Alberto Freitas

Bibliographic record

VenueJMIR Research Protocols · 2023
Typereview
Languageen
FieldComputer Science
TopicPrivacy-Preserving Technologies in Data
Canadian institutionsnot available
FundersEuropean Regional Development FundITEA
KeywordsComputer scienceImplementationHealth careData scienceMultinational corporationBig dataData qualityPsychological interventionInformation privacyRisk analysis (engineering)BusinessComputer securityMedicineData miningEconomicsNursingMarketing

Abstract

fetched live from OpenAlex

BACKGROUND: Considering the soaring health-related costs directed toward a growing, aging, and comorbid population, the health sector needs effective data-driven interventions while managing rising care costs. While health interventions using data mining have become more robust and adopted, they often demand high-quality big data. However, growing privacy concerns have hindered large-scale data sharing. In parallel, recently introduced legal instruments require complex implementations, especially when it comes to biomedical data. New privacy-preserving technologies, such as decentralized learning, make it possible to create health models without mobilizing data sets by using distributed computation principles. Several multinational partnerships, including a recent agreement between the United States and the European Union, are adopting these techniques for next-generation data science. While these approaches are promising, there is no clear and robust evidence synthesis of health care applications. OBJECTIVE: The main aim is to compare the performance among health data models (eg, automated diagnosis and mortality prediction) developed using decentralized learning approaches (eg, federated and blockchain) to those using centralized or local methods. Secondary aims are comparing the privacy compromise and resource use among model architectures. METHODS: We will conduct a systematic review using the first-ever registered research protocol for this topic following a robust search methodology, including several biomedical and computational databases. This work will compare health data models differing in development architecture, grouping them according to their clinical applications. For reporting purposes, a PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 flow diagram will be presented. CHARMS (Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies)-based forms will be used for data extraction and to assess the risk of bias, alongside PROBAST (Prediction Model Risk of Bias Assessment Tool). All effect measures in the original studies will be reported. RESULTS: The queries and data extractions are expected to start on February 28, 2023, and end by July 31, 2023. The research protocol was registered with PROSPERO, under the number 393126, on February 3, 2023. With this protocol, we detail how we will conduct the systematic review. With that study, we aim to summarize the progress and findings from state-of-the-art decentralized learning models in health care in comparison to their local and centralized counterparts. Results are expected to clarify the consensuses and heterogeneities reported and help guide the research and development of new robust and sustainable applications to address the health data privacy problem, with applicability in real-world settings. CONCLUSIONS: We expect to clearly present the status quo of these privacy-preserving technologies in health care. With this robust synthesis of the currently available scientific evidence, the review will inform health technology assessment and evidence-based decisions, from health professionals, data scientists, and policy makers alike. Importantly, it should also guide the development and application of new tools in service of patients' privacy and future research. TRIAL REGISTRATION: PROSPERO 393126; https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=393126. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/45823.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.045
metaresearch head score (Gemma)0.120
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Scholarly communication, Open science, Research integrity
Consensus categoriesMetaresearch, Open science
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: none
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.281
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0450.120
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0080.001
Bibliometrics0.0010.004
Science and technology studies0.0010.000
Scholarly communication0.0020.002
Open science0.1170.163
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.862
GPT teacher head0.728
Teacher spread0.134 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designSystematic review
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research ProtocolsSame topicPrivacy-Preserving Technologies in DataFrench-language works237,207