Accuracy of Computer-Aided Detection of Occupational Lung Disease: Silicosis and Pulmonary Tuberculosis in Ex-Miners from the South African Gold Mines
Bibliographic record
Abstract
BACKGROUND: Computer-aided detection (CAD) of pulmonary tuberculosis (TB) and silicosis among ex-miners from the South African gold mines has the potential to ease the backlog of lung examinations in clinical screening and medical adjudication for miners' compensation. This study aimed to determine whether CAD systems developed to date primarily for TB were able to identify TB (without distinction between prior and active disease) and silicosis (or "other abnormality") in this population. METHODS: A total of 501 chest X-rays (CXRs) from a screening programme were submitted to two commercial CAD systems for detection of "any abnormality", TB (any) and silicosis. The outcomes were tested against the readings of occupational medicine specialists with experience in reading miners' CXRs. Accuracy of CAD against the readers was calculated as the area under the curve (AUC) of the receiver operating characteristic (ROC) curve. Sensitivity and specificity were derived using a threshold requiring at least 90% sensitivity. RESULTS: One system was able to detect silicosis and/or TB with high AUCs (>0.85) against both readers, and specificity > 70% in most of the comparisons. The other system was able to detect "any abnormality" and TB with high AUCs, but with specificity < 70%. CONCLUSION: CAD systems have the potential to come close to expert readers in the identification of TB and silicosis in this population. The findings underscore the need for CAD systems to be developed and validated in specific use-case settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".