MétaCan
Menu
← Back to cohort
Record W4402966677 · doi:10.2196/63583

Harnessing Big Heterogeneous Data to Evaluate the Potential Impact of HIV Responses Among Key Populations in Sub-Saharan Africa: Protocol for the Boloka Data Repository Initiative

2024· article· en· W4402966677 on OpenAlexvenueno aff
Nancy Phaswana‐Mafuya, Edith Phalane, Amrita Rao, Kalai Willis, Katherine B. Rucinski, Karen Alida Voet, Amal Abdulrahman, Claris Siyamayambo, Betty Sebati, Mohlago Ablonia Seloka, Musa Jaiteh, Lerato Lucia Olifant, Katharine Journeay, Haley Sisel, Xiaoming Li, Bankole Olatosi, Neşet Hikmet, Prashant Duhoon, Francois Wolmarans, Yegnanew A. Shiferaw, Lifutso Motsieloa, Mashudu Rampilo, Stefan Baral

Bibliographic record

VenueJMIR Research Protocols · 2024
Typearticle
Languageen
FieldMedicine
TopicHIV/AIDS Research and Interventions
Canadian institutionsnot available
FundersNational Institute on Minority Health and Health DisparitiesNational Institute of Allergy and Infectious DiseasesNational Institute of Mental Health
KeywordsPreprintProtocol (science)Human immunodeficiency virus (HIV)Key (lock)Data scienceBig dataMedicineComputer scienceEnvironmental healthWorld Wide WebAlternative medicineFamily medicineComputer securityData mining

Abstract

fetched live from OpenAlex

BACKGROUND: In South Africa, there is no centralized HIV surveillance system where key populations (KPs) data, including gay men and other men who have sex with men, female sex workers, transgender persons, people who use drugs, and incarcerated persons, are stored in South Africa despite being on higher risk of HIV acquisition and transmission than the general population. Data on KPs are being collected on a smaller scale by numerous stakeholders and managed in silos. There exists an opportunity to harness a variety of data, such as empirical, contextual, observational, and programmatic data, for evaluating the potential impact of HIV responses among KPs in South Africa. OBJECTIVE: This study aimed to leverage and harness big heterogeneous data on HIV among KPs and harmonize and analyze it to inform a targeted HIV response for greater impact in Sub-Saharan Africa. METHODS: The Boloka data repository initiative has 5 stages. There will be engagement of a wide range of stakeholders to facilitate the acquisition of data (stage 1). Through these engagements, different data types will be collated (stage 2). The data will be filtered and screened to enable high-quality analyses (stage 3). The collated data will be stored in the Boloka data repository (stage 4). The Boloka data repository will be made accessible to stakeholders and authorized users (stage 5). RESULTS: The protocol was funded by the South African Medical Research Council following external peer reviews (December 2022). The study received initial ethics approval (May 2022), renewal (June 2023), and amendment (July 2024) from the University of Johannesburg (UJ) Research Ethics Committee. The research team has been recruited, onboarded, and received non-web-based internet ethics training (January 2023). A list of current and potential data partners has been compiled (January 2023 to date). Data sharing or user agreements have been signed with several data partners (August 2023 to date). Survey and routine data have been and are being secured (January 5, 2023). In (September 2024) we received Ghana Men Study data. The data transfer agreement between the Pan African Centre for Epidemics Research and the Perinatal HIV Research Unit was finalized (October 2024), and we are anticipating receiving data by (December 2024). In total, 7 abstracts are underway, with 1 abstract completed the analysis and expected to submit the full article to the peer-reviewed journal in early January 2024. As of March 2025, we expect to submit the remaining 6 full articles. CONCLUSIONS: A truly "complete" data infrastructure that systematically and rigorously integrates diverse data for KPs will not only improve our understanding of local epidemics but will also improve HIV interventions and policies. Furthermore, it will inform future research directions and become an incredible institutional mechanism for epidemiological and public health training in South Africa and Sub-Saharan Africa. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/63583.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.188
metaresearch head score (Gemma)0.226
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.812
Threshold uncertainty score0.993

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1880.226
Meta-epidemiology (narrow)0.0020.004
Meta-epidemiology (broad)0.0020.005
Bibliometrics0.0060.006
Science and technology studies0.0060.006
Scholarly communication0.0080.006
Open science0.0060.011
Research integrity0.0060.009
Insufficient payload (model declined to judge)0.0890.021

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.602
GPT teacher head0.619
Teacher spread0.016 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research Protocols→Same topicHIV/AIDS Research and Interventions→French-language works237,207→