Diagnostic Test Accuracy in Childhood Pulmonary Tuberculosis: A Bayesian Latent Class Analysis
Bibliographic record
Abstract
Evaluation of tests for the diagnosis of childhood pulmonary tuberculosis (CPTB) is complicated by the absence of an accurate reference test. We present a Bayesian latent class analysis in which we evaluated the accuracy of 5 diagnostic tests for CPTB. We used data from a study of 749 hospitalized South African children suspected to have CPTB from 2009 to 2014. The following tests were used: mycobacterial culture, smear microscopy, Xpert MTB/RIF (Cepheid Inc.), tuberculin skin test (TST), and chest radiography. We estimated the prevalence of CPTB to be 27% (95% credible interval (CrI): 21, 35). The sensitivities of culture, Xpert, and smear microscopy were estimated to be 60% (95% CrI: 46, 76), 49% (95% CrI: 38, 62), and 22% (95% CrI: 16, 30), respectively; specificities of these tests were estimated in accordance with prior information and were close to 100%. Chest radiography was estimated to have a sensitivity of 64% (95% CrI: 55, 73) and a specificity of 78% (95% CrI: 73, 83). Sensitivity of the TST was estimated to be 75% (95% CrI: 61, 84), and it decreased substantially among children who were malnourished and infected with human immunodeficiency virus (56%). The specificity of the TST was 69% (95% CrI: 63%, 76%). Furthermore, it was estimated that 46% (95% CrI: 42, 49) of CPTB-negative cases and 93% (95% CrI: 82; 98) of CPTB-positive cases received antituberculosis treatment, which indicates substantial overtreatment and limited undertreatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.206 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".