Bibliographic record
Abstract
We feel that Baluja and her coauthors1 missed an opportunity to address a paradox. Rates of smoking among male Chinese American youths are remarkably low,2 yet smoking rates among Chinese males living in China are said to exceed 60%.3 Baluja and associates reported the smoking prevalence rate for Chinese American immigrant males to be around 13%, which was well below the corresponding US male smoking prevalence rate of 24%.1 From these results, we infer that selection pressures involved in the immigration of Chinese to the United States favor Chinese who do not smoke. Indeed, immigrants generally exhibit healthier lifestyle practices and experience lower mortality rates than demographically similar US natives.4 We also feel that the authors could have been more proactive in warning readers about 2 additional limitations of their data. First, no questions were asked about tobacco products rarely used by Americans but commonly used overseas, such as bidis or paan. The low rates of smoking reported for US immigrants from India compared with US immigrants from other Asian countries might lead the naive reader to assume that overall tobacco use among immigrant Indians was low. Such an inference may be incorrect. Recent worldwide comparisons suggested that Indian women had the highest world incidence of oral cancer, attributable to their habit of chewing betel nut mixed with tobacco.5 Babu estimated that 90% of oral cancers in India were tobacco-related.6 In fact, estimates for 20007 suggested that oral cancers were the most important cause of cancer death in Indian males and the no. 2 cause of non–reproductive organ cancer death in Indian females. Comparisons between countries in use of just 1 tobacco product may hence suggest misleadingly low rates of overall tobacco use for some countries and their emigrants. A second additional limitation of the data was the implicit assumption that “country of origin” was a highly sensitive way to capture US immigrants of Asian national origin. In fact, US immigration quotas compelled many would-be immigrants to sojourn in Canada, Mexico, or Europe before migrating to the United States. Historic diasporas, such as the emigration of Indian nationals to British commonwealth countries, have also contributed to 2-step migrations to the United States. It is likely, then, that some US immigrants of Indian national origin were not counted in the results reported by Baluja and her associates because they listed a country other than India as their most recent country of origin.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".