Bibliographic record
Abstract
We conducted a molecular epidemiologic study of tuberculosis (TB) transmission during the years 1996--98 on the Island of Montreal. By combining public health data on the 528 reported cases with IS6110 DNA fingerprints for 430, we detected an overall low frequency of transmission manifesting as secondary cases of active disease. We also identified an important sub-group of TB patients who harbour isolates with matching patterns. Depending on the matching criterion used we attributed between 6% (95% CI: 4, 9%) and 22% (95% CI: 18, 27%) of TB cases to recent transmission; the vast majority of active TB disease reflects infection acquired at an earlier time and/or a different place. However, Haitian-born TB patients yielded a disproportionately high frequency of isolates belonging to matching "clusters" (21%; 95% CI: 13, 32%), while, other foreign-born patients have disproportionately low numbers of clustered isolates (5%; 95% CI: 3, 9%). The classical interpretation of such results is that there is more ongoing transmission within this immigrant sub-group. We explored an alternative hypothesis: that M. tuberculosis isolates from Haitian-born patients demonstrate reduced genetic diversity reflecting TB transmission patterns in their previously isolated country of origin---hence that a bacterial founder effect accounts for the higher frequency of matching fingerprints. Using a recently introduced measure of fingerprint similarity, genetic distance, we assessed the extent of pattern diversity. The median nearest genetic distance (NGD) was 130 months (inter-quartile range (IQR): 98--201 months) among the 47 distinct isolates from Haitian-born patients; among the non-Haitian foreign-born, the median NGD for the 191 distinct isolates was 128 months (IQR: 103--170 months). Hence the overall genetic heterogeneity of M. tuberculosis organisms among Haitian-born Montrealers was as great as that among a group of patients born in 70 other countries. Local transmission among the Haitian-born remains the most likely scenario. We demonstrated that a continuous measure, such as genetic distance, may also permit researchers to address a challenge to the interpretation of M. tuberculosis molecular typing results: how to determine whether highly similar, non-identical fingerprint patterns in fact reflect underlying "matches." The distribution of NGD for isolates initially classified as identical (10--27 months), similar (15--108 months) and unique (40--244 months) suggested a possible cut-point of 40 months. Use of this cut-point labelled 19% of isolates as "clustered", suggesting that 14% of Montreal TB cases reflected transmission during the study period.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".