Robust Tobacco Smoking Self-Report in two Cohorts of Vulnerable Pregnant Women and Adults
Bibliographic record
Abstract
Abstract Background Stigma associated with tobacco smoking, especially during pregnancy, may lead to underreporting and possible bias in studies relying on self-reported smoking data. Cotinine, a nicotine metabolite with a ∼20h half-life in blood, is often used as a biomarker of smoking. The objective of this study was to examine the concordance between self-reported smoking and plasma cotinine concentration among participants enrolled in two related cohorts of vulnerable individuals: human immunodeficiency virus (HIV)-positive and HIV-negative pregnant women enrolled in the CARMA-PREG cohort and HIV-positive and HIV-negative non-pregnant women and men enrolled in the CARMA-CORE cohort. Methods For HIV-positive (n=76) and negative (n=24) pregnant women, plasma cotinine was measured by ELISA in specimens collected during the third trimester, between 28 and 38 weeks of gestation. Plasma cotinine was also measured in HIV-positive (n=43) and negative (n=57) women and men enrolled in the CARMA-CORE cohort. Results Self-reported smokers were more likely to have low income (p<0.001) in both cohorts, and to deliver preterm (p=0.007) in CARMA-PREG. In the CARMA-PREG cohort, concordance between plasma cotinine was 95% for self-reported smoking, and 89% for self-reported non-smoking. In the CARMA-CORE cohort we observed similarly high concordances of 96% and 92% for self-reported smoking and non-smoking, respectively. In this sample, the odds of discordance between self-reported smoking status and cotinine levels were not significantly different between self-reported smokers and non-smokers, nor between pregnant women and others. Taken together, the overall concordance between plasma cotinine and self-reported data was 94% with a Cohen’s kappa coefficient of 0.860 among all participants. Conclusions Given the high proportion of vulnerable people in the CARMA-PREG and CARMA-CORE cohorts, our results may not be fully generalizable to the general population. However, they demonstrate that participant surveying in a non-judgemental context can lead to accurate and robust self-report data. Implications Reliable self-reported smoking data is necessary to account for smoking status in subsequent studies. Our results suggest that future studies should ensure that study participants feel sale to speak candidly to non-judgemental research staff to obtain reliable self-report data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".