TUBERCULOSIS GENOTYPING NETWORK Genotyping Analyses of Tuberculosis Cases in U.S.and Foreign-Born
Bibliographic record
Abstract
We used molecular genotyping to further understand the epidemiology and transmission patterns of tuberculosis (TB) in Massachusetts. The study population included 983 TB patients whose cases were verified by the Massachusetts Department of Public Health between July 1, 1996, and December 31, 2000, and for whom genotyping results and information on country of origin were available. Two hundred seventy-two (28%) of TB patients were in genetic clusters, and isolates from U.S-born were twice as likely to cluster as those of foreign-born (odds ratio [OR] 2.29, 95 % confidence interval [CI] 1.69 to 3.12). Our results suggest that restriction fragment length polymorphism analysis has limited capacity to differentiate TB strains when the isolate contains six or fewer copies of IS6110, even with spoligotyping. Clusters of TB patients with more than six copies of IS6110 were more likely to have epidemiologic connections than were clusters of TB patients with isolates with few copies of IS6110 (OR 8.01, 95%; CI 3.45 to 18.93). T he incidence of tuberculosis (TB) in the United States is closely linked to the global TB epidemic (1). In 2000, 46 % of all reported TB cases in the United States occurred among persons not born in the United States (foreign-born), and 20 states reported that>50 % of TB cases occurred among the foreign-born (2). In Massachusetts, 202 (71%) of 285 cases reported were among foreign-born persons (from 41 different countries). Being born outside the United States is the primary risk factor for being reported with TB in Massachusetts (3). The distribution of places of birth among TB patients reported in Massachusetts has changed greatly over the past 3 decades, reflecting changes in populations immigrating to Massachusetts. As late as 1970, 80 % of foreign immigrants in Massachusetts were from Europe or Canada; only 5 % of the immigrants were from Asia, and less than 3 % were from Central and South America combined and Africa (4). Since 1970, the proportion of immigrants to Massachusetts from Europe has declined, and the proportion of those from Asia, the Caribbean
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".