Determinants of Equitable Data Governance for African, Caribbean, and Black Communities in Health Research in High-Income Countries: Protocol for a Scoping Review
Bibliographic record
Abstract
BACKGROUND: African, Caribbean, and Black (ACB) communities in high-income countries continue to experience persistent health inequities, driven by systemic anti-Black racism, socioeconomic disadvantage, and exclusion from health decision-making. Historically, data have been extracted from ACB communities without transparency, accountability, or community ownership. These inequitable practices have produced data systems that reinforce harm rather than promote equity. Equitable data governance, which promotes community ownership over data collection, access, and use, is increasingly recognized as a critical but underresearched determinant of health equity. OBJECTIVE: This protocol outlines the methodology of a scoping review to identify and synthesize evidence on the determinants of equitable data governance in health research involving ACB communities in high-income countries. METHODS: The review follows the 6-stage Arksey and O'Malley methodological framework, supplemented with updated guidance from the Joanna Briggs Institute. The searches were conducted in the Ovid MEDLINE, Ovid Embase, EBSCO CINAHL, APA PsycInfo, and Scopus databases. Peer-reviewed articles are considered, with no limits placed on study design, publication type, or date. Multiple reviewers will independently extract data by using a standardized form. A 3-phase thematic mapping process, conceptually informed by critical race theory, intersectionality, and community-based participatory research principles, will be conducted to analyze the data, generate themes, and interpret findings. RESULTS: The final comprehensive database searches were completed on December 17, 2024. The search strategy targeted literature on data management, governance, sharing, security, and ethical principles in relation to ACB populations in high-income countries. A total of 4365 records were screened at the title and abstract level, after deduplication, of which 247 studies were deemed potentially relevant and advanced to full-text screening. Following full-text screening and reference list searching a total of 15 articles were deemed eligible for analysis. The data extraction stage is scheduled to overlap and occur between November 2025 and February 2026. The thematic mapping and stakeholder consultations processes are scheduled between December 2025 and February 2026. The final review and manuscript submission are expected by March 2026, with dissemination activities planned for mid-2026. CONCLUSIONS: This review will synthesize existing information on key pillars, barriers, facilitators, promising data governance policies and practices, and recommendations relevant to ACB communities. The findings may inform the expansion of Ontario's Engagement, Governance, Access, and Protection guidelines and support tailored research and national data governance frameworks. The review is expected to contribute to policy, research, and community-led data initiatives. Dissemination will occur through academic publications, conferences, and community-based knowledge-sharing events. As the review relies solely on publicly available data, ethics approval is not required. TRIAL REGISTRATION: OSF Registries 10.17605/OSF.IO/Z82AY; https://osf.io/z82ay. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/82403.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.147 | 0.151 |
| Meta-epidemiology (narrow) | 0.006 | 0.006 |
| Meta-epidemiology (broad) | 0.012 | 0.018 |
| Bibliometrics | 0.022 | 0.021 |
| Science and technology studies | 0.007 | 0.007 |
| Scholarly communication | 0.010 | 0.011 |
| Open science | 0.007 | 0.010 |
| Research integrity | 0.011 | 0.009 |
| Insufficient payload (model declined to judge) | 0.059 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".