From community as data providers to community as data users: developing a community-led research platform using program data in HIV/STI Program Science in Kenya
Bibliographic record
Abstract
Abstract Background Community-based organizations (CBOs) are critical in providing trusted and targeted HIV/STI services to gay, bisexual, and other men who have sex with men (GBMSM). Despite significant strides in CBOs’ involvement in HIV/STI research, there remain gaps in meaningful engagement, especially in quantitative research. This paper explores the development of HEKA, a community-led research platform where community-based organizations build capacity and leverage routinely collected program data to design research that aims to improve HIV/STI programs. We share a collective reflection on the lessons learned in the process, the challenges that emerged, and recommendations for facilitating community-based program science. Methodology Through a collaborative process, seven CBOs serving GBMSM in Kenya created the HEKA Research Initiative and designed a framework of collaboration, through which we assessed the technical gaps in quantitative research among staff, applied for funding, co-designed capacity-building workshops with academic partners, and developed a research agenda. We established a monthly meeting frequency and through collective reflection, documented the lessons and challenges in the process. Outcomes With our successful grant, we organized an in-person workshop on quantitative research methods and R programming. The team identified research questions and completed data cleaning/harmonization of program data. HEKA was successful because we emphasized a co-leadership framework (research direction evolved through shared/delegated leadership), and peer-to-peer mentorship. Major challenges included: obtaining sustained funding for engagement; ensuring the learning pace allows all individuals to be on the same page; confronting the socio-political climate; long commutes between counties for in-person meetings; and the limitation in using Excel files as primary tools for data capture. Conclusions HEKA demonstrates the potential for community-based and led research in the HIV/STI field. The model we present can serve as a blueprint for other community-based organizations aiming to lead collaborative or independent research and build capacity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.161 | 0.098 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.022 | 0.010 |
| Scholarly communication | 0.013 | 0.014 |
| Open science | 0.005 | 0.034 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".