A tutorial on a marginal structural modeling approach to mediation analysis in occupational health research: Investigating education, employment quality, and mortality
Bibliographic record
Abstract
Abstract Life expectancy inequities between more‐ and less‐educated groups have grown by 1 to 2 years over the last several decades in the United States. Simultaneously, employment conditions for many workers have deteriorated. Researchers hypothesize that these adverse conditions mediate educational inequities in mortality. However, methodological barriers have impeded research on the role of employment conditions and other hazards as mediating factors in health inequities. Indeed, traditional mediation analysis methods are often biased in occupational health settings, including in those with exposure‐mediator interactions and mediator‐outcome confounders that are caused by exposure. In this paper, we outline—and provide code for—a marginal structural modeling (MSM) approach for estimating total effects and controlled direct effects originally proposed elsewhere, which can be applied to common mediation analysis settings in occupational health research. As an example, we apply our approach to assess the extent to which disparities in employment quality (EQ)—a multidimensional construct characterizing the terms and conditions of the worker‐employer relationship—explained educational inequities in mortality in a 1999–2015 US Panel Study of Income Dynamics sample of workers with mortality follow‐up through 2017. Under certain strong assumptions described in the text, our estimates suggest that over 70% of the educational inequity in mortality would have been eliminated if EQ had been at the 80th percentile (100th = best) across exposure groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.048 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.003 | 0.007 |
| Insufficient payload (model declined to judge) | 0.025 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".