Bibliographic record
Abstract
To the Editor: We read with interest the recent article by Cefalu and Dominici1 regarding a linked statistical model for the assessment of spatially-dependent exposures and health. Our concerns center around the suggestion that this model can be applied to any epidemiologic design, the authors’ definition of confounding, that no personal risk factors were included in their health model, and that area-wide predictors for the exposure model are also confounding factors in a health model. As the health model used normally distributed errors, we assume that the authors were referring to a cross-sectional study of continuous outcomes. Except possibly for an ecologic study, we know of no design in which personal risk factors would not be included as potential confounding variables. In terms of what is a confounder, the authors stated “We refer to confounding bias as the bias in the health-effect estimate from a health-effects regression model that fails to control for any confounding….” This is not the accepted definition of confounding: quoting Breslow and Day from 1980,2 “Confounding is intimately connected to the concept of causality. …if some exposure E is associated with disease status, then the incidence of the disease varies among the strata defined by different level of E. If these differences in incidence are caused (partially) by some factor C, then we say that C has (partially) confounded the association between E and the disease.” We note that any noncausal variable could be associated with exposure and health, and these variables should not be included in a model to control for bias unless they are surrogates of causal processes. In our experience, there are very few area-level variables included in exposure models that are true causal variables for health outcomes. For example, in our land-use regression model of NO23 that was used in case–control studies on breast cancer,4 the predictors of traffic-related exposure included population density, counts of traffic, and distance to roads. None of these variables are causal risk factors for these cancers. Could any of these variables represent some complex causal process that can affect the incidence of these cancers? Possibly, but one would have to postulate the purported mechanism. For example, green space may lower pollution, so that the effects of greenspace on health could be due in part to lower levels of exposure to air pollution: this is a measurement issue and is not confounding. While one might contend that these contextual variables represent causal exposures, we suggest that these variables be modeled directly. Mark S. Goldberg Division of Clinical Epidemiology Department of Medicine McGill University Health Center – RVH Montreal, QC, Canada [email protected] Paul Villeneuve Department of Health Sciences Carleton University Ottawa, ON, Canada Daniel Crouse Department of Sociology University of New Brunswick Fredericton, NB, Canada
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.005 | 0.004 |
| Insufficient payload (model declined to judge) | 0.008 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".