Scoring systems using chest radiographic features for the diagnosis of pulmonary tuberculosis in adults: a systematic review
Bibliographic record
Abstract
Chest radiography for the diagnosis of active pulmonary tuberculosis (PTB) is limited by poor specificity and reader inconsistency. Scoring systems have been employed successfully for improving the performance of chest radiography for various pulmonary diseases. We conducted a systematic review to assess the diagnostic accuracy and reproducibility of scoring systems for PTB. We searched multiple databases for studies that evaluated the accuracy and reproducibility of chest radiograph scoring systems for PTB. We summarised results for specific radiographic features and scoring systems associated with PTB. Where appropriate, we estimated pooled performance of similar studies using a random effects model. 13 studies were included in the review, nine of which were in low tuberculosis (TB) burden settings. No scoring system was based solely on radiographic findings. All studies used systems with various combinations of clinical and radiological features. 11 studies involved scoring systems that were used for making decisions concerning hospital respiratory isolation. None of the included studies reported data on intra- or inter-reporter reproducibility. Upper lobe infiltrates (pooled diagnostic OR 3.57, 95% CI 2.38-5.37, five studies) and cavities (diagnostic OR range 1.97-25.66, three studies) were significantly associated with PTB. Sensitivities of the scoring systems were high (median 96%, IQR 93-98%), but specificities were low (median 46%, IQR 35-50%). Chest radiograph scoring systems appear useful in ruling out PTB in hospitals, but their low specificity precludes ruling in PTB. There is a need to develop accurate scoring systems for people living with HIV and for outpatient settings, especially in high TB burden settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".