Developing and Validating an Inclusive and Cost-Effective Prediction Algorithm for Survival and Death Among People Living With HIV in Sub-Saharan Africa: Protocol for a Meta-Analysis and Case-Control and Cost-Effectiveness Study
Bibliographic record
Abstract
BACKGROUND: Premature death in people with HIV in sub-Saharan Africa (SSA) is highly preventable. However, the lack of inclusive, cost-effective prognostic tools remains challenging. Most prognostic tools have been developed in high-income economies. The distinct cultural dynamics in HIV-related death epidemiology makes them unsuitable for the region. Additionally, the models lack systematic stratification of death determinants based on clinical relevance, and some included factors are too expensive for people with HIV in SSA. OBJECTIVE: We aimed to create a tailored predictive model that considers the unique context of SSA, including cultural dynamics, cost-effectiveness, and clinical relevance. METHODS: This is a 2-phase study. In the development phase, we will use a combination of evidence synthesis, namely meta-analysis, application epidemiology, biostatistical, and economic paradigms, to develop a prognostic model for people living with HIV in SSA. The Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) protocol will be followed in the structuring of the meta-analysis. From their creation to the present, we will search African journals (Sabinet) and the PubMed, Scopus, MEDLINE, Academic Search Complete, Directory of Open Access Repository, Cochrane Library, Web of Science, EMBASE, and Cumulative Index for Nursing and Allied Health Literature databases. Only cohort studies with moderate to high quality will be included. The primary outcome variables include the predictors of HIV-related death and their corresponding effect sizes (adjusted relative risk). A random-effect meta-analysis model will be used to synthesize the unbiased estimate of risk (relative risk) per predictor. Epidemiological metrics such as risk responsiveness, geotemporal trend, risk weight (Rw), clinical minimum important difference (CMID), predictors interaction density (PID), critical risk points, and potential cost implication will be computed. A combination of Rw and CMID will be used for risk stratification. The model's constituent items will be selected based on the combination of Rw, CMID, PID and cost implication. In the validation phase, we will apply the emergent model to classify participants using a secondary data obtained from a cohort of people living with HIV in East and West Africa, with outcomes including sensitivity, specificity, calibration, and area under the receiver operating characteristic curve (AUC). RESULTS: The study is projected to commence in October 2025 and end in September 2026. The expected result will be published in November 2026. The result will be presented using narrative and quantitative synthesis. Indices of causality namely as strength of association, temporality, consistency, biological gradient, and specificity of the predictor-outcome association will be presented in a tabular format. TheAUC will be used to decide the optimal critical risk point for the emergent predictive algorithm. CONCLUSIONS: Effective prognostication coupled with intense monitoring and evaluation, and prioritizing of therapeutic targets could positively turn around the fate of millions of people living with HIV at risk of premature death in SSA. TRIAL REGISTRATION: PROSPERO CRD42023430437; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023430437. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/63783.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.112 | 0.184 |
| Meta-epidemiology (narrow) | 0.005 | 0.003 |
| Meta-epidemiology (broad) | 0.014 | 0.033 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.007 | 0.006 |
| Insufficient payload (model declined to judge) | 0.039 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".