Overview of the Reanalysis of the Harvard Six Cities Study and American Cancer Society Study of Particulate Air Pollution and Mortality
Bibliographic record
Abstract
This article provides an overview of the Reanalysis Study of the Harvard Six Cities and the American Cancer Society (ACS) studies of particulate air pollution and mortality. The previous findings of the studies have been subject to debate. In response, a reanalysis team, comprised of Canadian and American researchers, was invited to participate in an independent reanalysis project to address the concerns. Phase I of the reanalysis involved the design of data audits to determine whether each study conformed to the consistency and accuracy of their data. Phase II of the reanalysis involved conducting a series of comprehensive analyses using alternative statistical methods. Alternative models were also used to identify covariates that may confound or modify the association of particulate air pollution as well as identify sensitive population subgroups. The audit demonstrated that the data in the original analyses were of high quality, as were the risk estimates reported by the original investigators. The sensitivity analysis illustrated that the mortality risk estimates reported in both studies were found to be robust against alternative Cox models. Detailed investigation of the covariate effects found a significant modifying effect of education and a relative risk of mortality associated with fine particles and declining education levels. The study team applied spatial analytic methods to the ACS data, resulting in various levels of spatial autocorrelations supporting the reported association for fine particles mortality of the original investigators as well as demonstrating a significant association between sulfur dioxide and mortality. Collectively, our reanalysis suggest that mortality may be attributable to more than one component of the complex mixture of ambient air pollutants for U.S. urban areas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".