Defining Care Patterns and Outcomes Among Persons Living with HIV in Washington, DC: Linkage of Clinical Cohort and Surveillance Data
Bibliographic record
Abstract
BACKGROUND: Triangulation of data from multiple sources such as clinical cohort and surveillance data can help improve our ability to describe care patterns, service utilization, comorbidities, and ultimately measure and monitor clinical outcomes among persons living with HIV infection. OBJECTIVES: The objective of this study was to determine whether linkage of clinical cohort data and routinely collected HIV surveillance data would enhance the completeness and accuracy of each database and improve the understanding of care patterns and clinical outcomes. METHODS: We linked data from the District of Columbia (DC) Cohort, a large HIV observational clinical cohort, with Washington, DC, Department of Health (DOH) surveillance data between January 2011 and June 2015. We determined percent concordance between select variables in the pre- and postlinked databases using kappa test statistics. We compared retention in care (RIC), viral suppression (VS), sexually transmitted diseases (STDs), and non-HIV comorbid conditions (eg, hypertension) and compared HIV clinic visit patterns determined using the prelinked database (DC Cohort) versus the postlinked database (DC Cohort + DOH) using chi-square testing. Additionally, we compared sociodemographic characteristics, RIC, and VS among participants receiving HIV care at ≥3 sites versus <3 sites using chi-square testing. RESULTS: Of the 6054 DC Cohort participants, 5521 (91.19%) were included in the postlinked database and enrolled at a single DC Cohort site. The majority of the participants was male, black, and had men who have sex with men (MSM) as their HIV risk factor. In the postlinked database, 619 STD diagnoses previously unknown to the DC Cohort were identified. Additionally, the proportion of participants with RIC was higher compared with the prelinked database (59.83%, 2678/4476 vs 64.95%, 2907/4476; P<.001) and the proportion with VS was lower (87.85%, 2277/2592 vs 85.15%, 2391/2808; P<.001). Almost a quarter of participants (23.06%, 1279/5521) were identified as receiving HIV care at ≥2 sites (postlinked database). The participants using ≥3 care sites were more likely to achieve RIC (80.7%, 234/290 vs 62.61%, 2197/3509) but less likely to achieve VS (72.3%, 154/213 vs 89.51%, 1869/2088). The participants using ≥3 care sites were more likely to have unstable housing (15.1%, 64/424 vs 8.96%, 380/4242), public insurance (86.1%, 365/424 vs 57.57%, 2442/4242), comorbid conditions (eg, hypertension) (37.7%, 160/424 vs 22.98%, 975/4242), and have acquired immunodeficiency syndrome (77.8%, 330/424 vs 61.20%, 2596/4242) (all P<.001). CONCLUSIONS: Linking surveillance and clinical data resulted in the improved completeness of each database and a larger volume of available data to evaluate HIV outcomes, allowing for refinement of HIV care continuum estimates. The postlinked database also highlighted important differences between participants who sought HIV care at multiple clinical sites. Our findings suggest that combined datasets can enhance evaluation of HIV-related outcomes across an entire metropolitan area. Future research will evaluate how to best utilize this information to improve outcomes in addition to monitoring them.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".