Population genomics of <i>Mycobacterium tuberculosis</i> in the Inuit
Bibliographic record
Abstract
Nunavik, Québec suffers from epidemic tuberculosis (TB), with an incidence 50-fold higher than the Canadian average. Molecular studies in this region have documented limited bacterial genetic diversity among Mycobacterium tuberculosis isolates, consistent with a founder strain and/or ongoing spread. We have used whole-genome sequencing on 163 M. tuberculosis isolates from 11 geographically isolated villages to provide a high-resolution portrait of bacterial genetic diversity in this setting. All isolates were lineage 4 (Euro-American), with two sublineages present (major, n = 153; minor, n = 10). Among major sublineage isolates, there was a median of 46 pairwise single-nucleotide polymorphisms (SNPs), and the most recent common ancestor (MRCA) was in the early 20th century. Pairs of isolates within a village had significantly fewer SNPs than pairs from different villages (median: 6 vs. 47, P < 0.00005), indicating that most transmission occurs within villages. There was an excess of nonsynonymous SNPs after the diversification of M. tuberculosis within Nunavik: The ratio of nonsynonymous to synonymous substitution rates (dN/dS) was 0.534 before the MRCA but 0.777 subsequently (P = 0.010). Nonsynonymous SNPs were detected across all gene categories, arguing against positive selection and toward genetic drift with relaxation of purifying selection. Supporting the latter possibility, 28 genes were partially or completely deleted since the MRCA, including genes previously reported to be essential for M. tuberculosis growth. Our findings indicate that the epidemiologic success of M. tuberculosis in this region is more likely due to an environment conducive to TB transmission than a particularly well-adapted strain.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".