Blood Transcriptomic Stratification of Short-term Risk in Contacts of Tuberculosis
Bibliographic record
Abstract
BACKGROUND: The highest risk of tuberculosis arises in the first few months after exposure. We reasoned that this risk reflects incipient disease among tuberculosis contacts. Blood transcriptional biomarkers of tuberculosis may predate clinical diagnosis, suggesting they offer improved sensitivity to detect subclinical incipient disease. Therefore, we sought to test the hypothesis that refined blood transcriptional biomarkers of active tuberculosis will improve stratification of short-term disease risk in tuberculosis contacts. METHODS: We combined analysis of previously published blood transcriptomic data with new data from a prospective human immunodeficiency virus (HIV)-negative UK cohort of 333 tuberculosis contacts. We used stability selection as an alternative computational approach to identify an optimal signature for short-term risk of active tuberculosis and evaluated its predictive value in independent cohorts. RESULTS: In a previously published HIV-negative South African case-control study of patients with asymptomatic Mycobacterium tuberculosis infection, a novel 3-gene transcriptional signature comprising BATF2, GBP5, and SCARF1 achieved a positive predictive value (PPV) of 23% for progression to active tuberculosis within 90 days. In a new UK cohort of 333 HIV-negative tuberculosis contacts with a median follow-up of 346 days, this signature achieved a PPV of 50% (95% confidence interval [CI], 15.7-84.3) and negative predictive value of 99.3% (95% CI, 97.5-99.9). By comparison, peripheral blood interferon gamma release assays in the same cohort achieved a PPV of 5.6% (95% CI, 2.1-11.8). CONCLUSIONS: This blood transcriptional signature provides unprecedented opportunities to target therapy among tuberculosis contacts with greatest risk of incident disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".