Tuberculosis screening among ambulatory people living with HIV: a systematic review and individual participant data meta-analysis
Bibliographic record
Abstract
BACKGROUND: The WHO-recommended tuberculosis screening and diagnostic algorithm in ambulatory people living with HIV is a four-symptom screen (known as the WHO-recommended four symptom screen [W4SS]) followed by a WHO-recommended molecular rapid diagnostic test (eg Xpert MTB/RIF [hereafter referred to as Xpert]) if W4SS is positive. To inform updated WHO guidelines, we aimed to assess the diagnostic accuracy of alternative screening tests and strategies for tuberculosis in this population. METHODS: In this systematic review and individual participant data meta-analysis, we updated a search of PubMed (MEDLINE), Embase, the Cochrane Library, and conference abstracts for publications from Jan 1, 2011, to March 12, 2018, done in a previous systematic review to include the period up to Aug 2, 2019. We screened the reference lists of identified pieces and contacted experts in the field. We included prospective cross-sectional, observational studies and randomised trials among adult and adolescent (age ≥10 years) ambulatory people living with HIV, irrespective of signs and symptoms of tuberculosis. We extracted study-level data using a standardised data extraction form, and we requested individual participant data from study authors. We aimed to compare the W4SS with alternative screening tests and strategies and the WHO-recommended algorithm (ie, W4SS followed by Xpert) with Xpert for all in terms of diagnostic accuracy (sensitivity and specificity), overall and in key subgroups (eg, by antiretroviral therapy [ART] status). The reference standard was culture. This study is registered with PROSPERO, CRD42020155895. FINDINGS: ), and lymphadenopathy had high specificities (80-90%) but low sensitivities (29-43%). The WHO-recommended algorithm had a sensitivity of 58% (50-66) and a specificity of 99% (98-100); Xpert for all had a sensitivity of 68% (57-76) and a specificity of 99% (98-99). In the one study that assessed both, the sensitivity of sputum Xpert Ultra was higher than sputum Xpert (73% [62-81] vs 57% [47-67]) and specificities were similar (98% [96-98] vs 99% [98-100]). Among outpatients on ART (4309 [99·1%] of 4347 people on ART), W4SS sensitivity was 53% (35-71) and specificity was 71% (51-85). In this population, a parallel strategy (two tests done at the same time) of W4SS with any chest x-ray abnormality had higher sensitivity (89% [70-97]) and lower specificity (33% [17-54]; n=2670) than W4SS alone; at a tuberculosis prevalence of 5%, this strategy would require 379 more rapid diagnostic tests per 1000 people living with HIV than W4SS but detect 18 more tuberculosis cases. Among outpatients not on ART (11 160 [71·8%] of 15 541 outpatients), W4SS sensitivity was 85% (76-91) and specificity was 37% (25-51). C-reactive protein (≥10 mg/L) alone had a similar sensitivity to (83% [79-86]), but higher specificity (67% [60-73]; n=3187) than, W4SS and a sequential strategy (both test positive) of W4SS then C-reactive protein (≥5 mg/L) had a similar sensitivity to (84% [75-90]), but higher specificity than (64% [57-71]; n=3187), W4SS alone; at 10% tuberculosis prevalence, these strategies would require 272 and 244 fewer rapid diagnostic tests per 1000 people living with HIV than W4SS but miss two and one more tuberculosis cases, respectively. INTERPRETATION: C-reactive protein reduces the need for further rapid diagnostic tests without compromising sensitivity and has been included in the updated WHO tuberculosis screening guidelines. However, C-reactive protein data were scarce for outpatients on ART, necessitating future research regarding the utility of C-reactive protein in this group. Chest x-ray can be useful in outpatients on ART when combined with W4SS. The WHO-recommended algorithm has suboptimal sensitivity; Xpert for all offers slight sensitivity gains and would have major resource implications. FUNDING: World Health Organization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".