Identifying Health Care Services Offered in the HIV Care Continuum via a Machine Learning–Based Topic Modeling Approach: Exploratory Literature Review
Bibliographic record
Abstract
Background: It remains unclear whether the existing health care services reflect the HIV care continuum, which underscores the need for integrated care beyond viral suppression. Objective: This study aimed to analyze the literature on health care services for people living with HIV to enhance the understanding of trends and knowledge structures. Methods: A literature review was conducted using BERTopic, an advanced machine learning-based topic modeling technique. We searched PubMed, CINAHL, EMBASE, and Cochrane databases for English-language studies published between 2013 and 2023. Analyses were performed twice: first, to gain a broad understanding of the literature, and second, to examine the specific details of health care services described. Results: Among the 11,269 articles screened, 204 studies met the inclusion criteria. Within the HIV care continuum, most studies focused on the treatment retention stage, while studies focusing on the long-term stage were limited. A broad literature analysis identified five key topics, with "ART adherence" emerging as the most prominent topic. A more comprehensive analysis of health care services within the literature revealed 7 topics, reflecting diverse delivery methods and content in providing health care services for people living with HIV. The predominant topic, "ART adherence and counseling," encompassed the largest number of studies, indicating the strongest emphasis in the field. Notably, the distribution of topics exhibited a distinct pattern: while health care service diversity was the highest in the earlier stages of the HIV care continuum, it became increasingly limited in the later stages. Conclusions: This study provides valuable insights into current HIV care services and highlights areas for future research and intervention. Despite the shift toward lifelong HIV management, existing literature remains heavily focused on medication treatment, overlooking the multifaceted health care needs of people living with HIV. Research disparities, particularly concerning vulnerable populations, underscore the need for more inclusive studies and tailored health care services. Efforts should be intensified to bridge these gaps, ensuring inclusive and equitable health care services across diverse populations and fostering interdisciplinary collaboration to meet the evolving needs of people living with HIV, thereby enhancing the HIV care continuum for all.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.064 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.008 |
| Bibliometrics | 0.061 | 0.049 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".