Evaluation of a Rapid Point of Care Test for Detecting Acute and Established HIV Infection, and Examining the Role of Study Quality on Diagnostic Accuracy: A Bayesian Meta-Analysis
Bibliographic record
Abstract
INTRODUCTION: Fourth generation (Ag/Ab combination) point of care HIV tests like the FDA-approved Determine HIV1/2 Ag/Ab Combo test offer the promise of timely detection of acute HIV infection, relevant in the context of HIV control. However, a synthesis of their performance has not yet been done. In this meta-analysis we not only assessed device performance but also evaluated the role of study quality on diagnostic accuracy. METHODS: Two independent reviewers searched seven databases, including conferences and bibliographies, and independently extracted data from 17 studies. Study quality was assessed with QUADAS-2. Data on sensitivity and specificity (overall, antigen, and antibody) were pooled using a Bayesian hierarchical random effects meta-analysis model. Subgroups were analyzed by blood samples (serum/plasma vs. whole blood) and study designs (case-control vs. cross-sectional). RESULTS: The overall specificity of the Determine Combo test was 99.1%, 95% credible interval (CrI) [97.3-99.8]. The overall pooled sensitivity for the device was at 88.5%, 95% [80.1-93.4]. When the components of the test were analyzed separately, the pooled specificities were 99.7%, 95% CrI [96.8-100] and 99.6%, 95% CrI [99.0-99.8], for the antigen and antibody components, respectively. Pooled sensitivity of the antibody component was 97.3%, 95% CrI [60.7-99.9], and pooled sensitivity for the antigen component was found to be 12.3%, 95% (CrI) [1.1-44.2]. No significant differences were found between subgroups by blood sample or study design. However, it was noted that many studies restricted their study sample to p24 antigen or RNA positive specimens, which may have led to underestimation of overall test performance. Detection bias, selection (spectrum) bias, incorporation bias, and verification bias impaired study quality. CONCLUSIONS: Although the specificity of all test components was high, antigenic sensitivity will merit from an improvement. Besides the accuracy of the device itself, study quality, also impacts the performance of the test. These factors must be kept in mind in future evaluations of an improved device, relevant for global scale up and implementation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.104 | 0.196 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.016 | 0.056 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".