MétaCan
Menu
Back to cohort
Record W4408139986 · doi:10.1093/ije/dyaf013

Protocol for improving equity in quantitative big data cleaning: lessons from longitudinal analysis of electronic health records from underrepresented and marginalized communities

2025· article· en· W4408139986 on OpenAlexaboutno aff
Zeruiah V. Buchanan, Scarlett E. Hopkins, Bert B. Boyer, Alison E. Fohner

Bibliographic record

VenueInternational Journal of Epidemiology · 2025
Typearticle
Languageen
FieldMedicine
TopicEthics in Clinical Research
Canadian institutionsnot available
FundersNational Institute on AgingNational Institutes of Health
KeywordsHealth equityProtocol (science)Equity (law)Underrepresented MinorityHealth recordsBig dataLongitudinal dataBusinessEnvironmental healthData scienceMedicineGerontologyPolitical scienceSociologyComputer scienceEconomic growthMedical educationPublic healthData miningDemographyHealth careNursingEconomicsAlternative medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Large biomedical datasets, including electronic health records (EHRs), are a significant source of epidemiologic data. To prepare an EHR for analysis, there are several data-cleaning approaches; here, we focus on data filtering. Common data-filtering methods employ rules that rely on data from socially constructed dominant populations but are inappropriate for marginalized populations, leading to the loss of valuable data and neglect of underrepresented communities. We propose a novel method based on a phenomenological framework that is more equitable and inclusive, leading to culturally responsive research and discoveries. METHODS: EHRs from the Yukon-Kuskokwim Health Corporation (YKHC) containing 1 262 035 records from 12 402 unique individuals from 2002 to 2012 were cleaned by using the proposed phenomenological (individual) and common (cohort) data-filtering approach. Within the phenomenological framework, we (i) excluded values that were undeniably biologically impossible for any population, (ii) excludes values that fell outside three standard deviations from the mean value for each individual person, and (iii) used two forms of imputation methods for stable quantitative and qualitative values at the individual level when data were missing. RESULTS: Compared with common data-filtering practices, the phenomenological approach retained more observations, participants, and a range of outcomes, allowing a truer representation of the priority population. In sensitivity analyses comparing the results of the raw data, the common approach implemented, and the phenomenological approach applied, we found that the phenomenological approach did not compromise the integrity of the results. CONCLUSION: The phenomenological approach to filtering big data presents an opportunity to better advocate for marginalized communities even when using large datasets that require automated rules for data filtering. Our method may empower researchers who are partnering with communities to embrace large datasets without compromising their commitment to community benefit and respect.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.497
metaresearch head score (Gemma)0.636
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Protocol · Consensus signal: none
Teacher disagreement score0.993
Threshold uncertainty score0.621

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4970.636
Meta-epidemiology (narrow)0.0020.003
Meta-epidemiology (broad)0.0030.005
Bibliometrics0.0070.010
Science and technology studies0.0090.010
Scholarly communication0.0070.008
Open science0.0070.011
Research integrity0.0100.015
Insufficient payload (model declined to judge)0.0190.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.860
GPT teacher head0.719
Teacher spread0.141 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of EpidemiologySame topicEthics in Clinical ResearchFrench-language works237,207