RE: "SMOKING AND PARKINSON'S DISEASE: USING PARENTAL SMOKING AS A PROXY TO EXPLORE CAUSALITY"
Bibliographic record
Abstract
We applaud O'Reilly et al. (1) for using directed acyclic graphs (DAGs) to convey structural assumptions in interpreting substantive research results. In their recent article, the authors showed an inverse association between parental smoking and offspring Parkinson's disease (PD) (1). They then stratified the results by offspring smoking status and reported no association between PD and parental smoking among the nonsmokers. As the authors noted, these results would be expected (notwithstanding chance, confounding, and other biases) if offspring smoking were an intermediate factor on a causal path from parental smoking to offspring PD (see the authors’ Figure 3), but not under the hypotheses that parental smoking has a direct effect on offspring PD and that prediagnostic PD in the offspring affects their smoking habits (“reverse prevention”; see the authors’ Figure 2). The DAGs O'Reilly et al. drew to depict these and other hypotheses raised additional questions, however, as clarifying devices often do. It seems unarguable that in all plausible structures in this context, parental smoking directly affects offspring smoking and offspring PD does not affect parental smoking. Those assumptions leave the 6 logically possible structures among the 3 core variables shown in our Table 1. One of them, in which parental smoking directly affects offspring PD but offspring smoking does not, might be too implausible to consider. We wonder, however, if the authors agree that it might be worthwhile to entertain the other 2 core structures in our Table 1. Each is composed of hypothetical effects the authors included in at least 1 of their DAGs. Logically Possible Causal Structures Among Parental Smoking, Offspring Smoking, and Offspring Parkinson's Disease, Given a Direct Effect of Parental Smoking on Offspring Smoking and No Effect of Offspring Parkinson's Disease on Parental Smoking Abbreviation: PD, Parkinson's disease. Logically Possible Causal Structures Among Parental Smoking, Offspring Smoking, and Offspring Parkinson's Disease, Given a Direct Effect of Parental Smoking on Offspring Smoking and No Effect of Offspring Parkinson's Disease on Parental Smoking Abbreviation: PD, Parkinson's disease. Each of O'Reilly et al.’s DAGs has a different array of covariates of concern as potential confounders (1). In Figures 1 and 3, the only such covariates are common causes of offspring smoking and offspring PD. Those covariates are absent from their Figure 5, which shows covariates affecting parental smoking and offspring PD. In Figure 4, every covariate that affects any 2 of the 3 core variables affects all of them. Figure 2 contains no covariates. It would be helpful for the authors to settle on a consistent set of covariate structures they consider reasonable and to show that set on the DAG for each plausible core structure (e.g., the 3 they have already considered plus the 2 suggested in our Table 1). Then the results of the data analysis could be interpreted in light of diagrams, each of which has its own set of confounding paths, but all of which are based on the same tenable configuration of covariate effects. Finally, it would be necessary to see the quantitative results from the authors’ analysis of parental smoking and offspring PD in all strata of offspring smoking, as opposed to an assertion that in 1 stratum there was “no association,” which can mean many things to many epidemiologists. For instance, the authors’ Figure 3 (given adequate control of the covariates depicted) predicts no association between parental smoking and offspring PD in every stratum of offspring smoking, not just the nonsmokers. In addition, without adequate control for the covariates in Figure 3, a null association could be observed between parental smoking and offspring PD within strata of offspring smoking, even in the presence of a strong direct effect of parental smoking on offspring PD (2, 3). We emphasize that these suggestions arose much more readily because the authors drew and reported on hypothetical DAGs in their paper (1), a practice we hope will become more widespread. When methodologists write generically about using DAGs, they understandably need to sidestep the whole “what's the right DAG” question. When researchers with substantive interests use DAGs, however, that question is the main and unavoidable one. Conflict of interest: none declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.091 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.005 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.069 | 0.093 |
| Insufficient payload (model declined to judge) | 0.008 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".