Development of a Simple Reliable Radiographic Scoring System to Aid the Diagnosis of Pulmonary Tuberculosis
Bibliographic record
Abstract
RATIONALE: Chest radiography is sometimes the only method available for investigating patients with possible pulmonary tuberculosis (PTB) with negative sputum smears. However, interpretation of chest radiographs in this context lacks specificity for PTB, is subjective and is neither standardized nor reproducible. Efforts to improve the interpretation of chest radiography are warranted. OBJECTIVES: To develop a scoring system to aid the diagnosis of PTB, using features recorded with the Chest Radiograph Reading and Recording System (CRRS). METHODS: Chest radiographs of outpatients with possible PTB, recruited over 3 years at clinics in South Africa were read by two independent readers using the CRRS method. Multivariate analysis was used to identify features significantly associated with culture-positive PTB. These were weighted and used to generate a score. RESULTS: 473 patients were included in the analysis. Large upper lobe opacities, cavities, unilateral pleural effusion and adenopathy were significantly associated with PTB, had high inter-reader reliability, and received 2, 2, 1 and 2 points, respectively in the final score. Using a cut-off of 2, scores below this threshold had a high negative predictive value (91.5%, 95%CI 87.1,94.7), but low positive predictive value (49.4%, 95%CI 42.9,55.9). Among the 382 TB suspects with negative sputum smears, 229 patients had scores <2; the score correctly ruled out active PTB in 214 of these patients (NPV 93.4%; 95%CI 89.4,96.3). The score had a suboptimal negative predictive value in HIV-infected patients (NPV 86.4, 95% CI 75,94). CONCLUSIONS: The proposed scoring system is simple, and reliably ruled out active PTB in smear-negative HIV-uninfected patients, thus potentially reducing the need for further tests in high burden settings. Validation studies are now required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".