The Bandim tuberculosis score: Reliability and comparison with the Karnofsky performance score
Bibliographic record
Abstract
BACKGROUND: This study was carried out in Guinea-Bissau's capital Bissau among inpatients and outpatients attending for tuberculosis (TB) treatment within the study area of the Bandim Health Project, a Health and Demographic Surveillance Site. Our aim was to assess the variability between 2 physicians in performing the Bandim tuberculosis score (TBscore), a clinical severity score for pulmonary TB (PTB), and to compare it to the Karnofsky performance score (KPS). METHOD: From December 2008 to July 2009 we assessed the TBscore and the KPS of 100 PTB patients at inclusion in the TB cohort and/or at 1 or more follow-up visits; 61 baseline and 130 follow-up double assessments were obtained. RESULTS: The inter-observer variability of the TBscore (5 symptoms and 6 clinical findings) varied from slight to almost perfect agreement. For the TBscore, all 3 severity classes (SC I-III) were observed, while the KPS only yielded 2 of its 3 possible classes. The grading of PTB patients into severity classes showed moderate agreement for both the TBscore (κ(w) = 0.52, 95% confidence interval 0.46-0.56) and the KPS (κ(w) = 0.49, 95% confidence interval 0.33-0.65). The intra-class correlation coefficient (ICC) was larger for the TBscore than for the KPS (0.822 vs 0.632). CONCLUSIONS: The Bandim TBscore had an acceptable inter-observer variability, seemed to be more disease-related, and performed better than the KPS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".