MétaCan
Menu
Back to cohort
Record W4382400983 · doi:10.2196/46471

Data Quality– and Utility-Compliant Anonymization of Common Data Model–Harmonized Electronic Health Record Data: Protocol for a Scoping Review

2023· review· en· W4382400983 on OpenAlexvenueno aff
Gaetan Kamdje Wabo, Fabian Praßer, Kerstin Gierend, Fabian Siegel, Thomas Ganslandt

Bibliographic record

VenueJMIR Research Protocols · 2023
Typereview
Languageen
FieldComputer Science
TopicPrivacy-Preserving Technologies in Data
Canadian institutionsnot available
FundersBundesministerium für Bildung und ForschungDeutsche Forschungsgemeinschaft
KeywordsComputer scienceData qualityContext (archaeology)Data miningData scienceHealth careInformation retrievalQuality (philosophy)Data anonymizationProtocol (science)Selection (genetic algorithm)Process (computing)Information privacyMedicineMetric (unit)Machine learningComputer security

Abstract

fetched live from OpenAlex

BACKGROUND: The anonymization of Common Data Model (CDM)-converted EHR data is essential to ensure the data privacy in the use of harmonized health care data. However, applying data anonymization techniques can significantly affect many properties of the resulting data sets and thus biases research results. Few studies have reviewed these applications with a reflection of approaches to manage data utility and quality concerns in the context of CDM-formatted health care data. OBJECTIVE: Our intended scoping review aims to identify and describe (1) how formal anonymization methods are carried out with CDM-converted health care data, (2) how data quality and utility concerns are considered, and (3) how the various CDMs differ in terms of their suitability for recording anonymized data. METHODS: The planned scoping review is based on the framework of Arksey and O'Malley. By using this, only articles published in English will be included. The retrieval of literature items should be based on a literature search string combining keywords related to data anonymization, CDM standards, and data quality assessment. The proposed literature search query should be validated by a librarian, accompanied by manual searches to include further informal sources. Eligible articles will first undergo a deduplication step, followed by the screening of titles. Second, a full-text reading will allow the 2 reviewers involved to reach the final decision about article selection, while a domain expert will support the resolution of citation selection conflicts. Additionally, key information will be extracted, categorized, summarized, and analyzed by using a proposed template into an iterative process. Tabular and graphical analyses should be addressed in alignment with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) checklist. We also performed some tentative searches on Web of Science for estimating the feasibility of reaching eligible articles. RESULTS: Tentative searches on Web of Science resulted in 507 nonduplicated matches, suggesting the availability of (potential) relevant articles. Further analysis and selection steps will allow us to derive a final literature set. Furthermore, the completion of this scoping review study is expected by the end of the fourth quarter of 2023. CONCLUSIONS: Outlining the approaches of applying formal anonymization methods on CDM-formatted health care data while taking into account data quality and utility concerns should provide useful insights to understand the existing approaches and future research direction based on identified gaps. This protocol describes a schedule to perform a scoping review, which should support the conduction of follow-up investigations. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/46471.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.177
metaresearch head score (Gemma)0.194
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.177
Threshold uncertainty score0.937

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1770.194
Meta-epidemiology (narrow)0.0050.005
Meta-epidemiology (broad)0.0100.012
Bibliometrics0.0190.017
Science and technology studies0.0060.007
Scholarly communication0.0100.009
Open science0.0060.009
Research integrity0.0110.008
Insufficient payload (model declined to judge)0.0600.019

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.901
GPT teacher head0.713
Teacher spread0.188 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSystematic review
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research ProtocolsSame topicPrivacy-Preserving Technologies in DataFrench-language works237,207