Reanalysis of the Harvard Six Cities Study, Part I: Validation and Replication
Bibliographic record
Abstract
Because the results of the Harvard Six Cities Study played a critical role in the establishment of the current U.S. ambient air quality objective for fine particles (PM(2.5)), the U.S. Environmental Protection Agency, industry, and nongovernmental organizations called for an independent reanalysis of this study to validate the original findings reported by Dockery and colleagues in the New England Journal of Medicine (vol. 329, pp. 1753-1759) in 1993. Validation of the original findings was accomplished by a detailed statistical audit and replication of original results. With the exception of occupational exposure to dust (14 discrepancies of 249 questionnaires located for evaluation) and fumes (15/249), date of death (2/250), and cause of death (2/250), the audit identified no discrepancies between the original questionnaires and death certificates in the audit sample and the analytic file used by the original investigators. The data quality audit identified a computer programming problem that had resulted in early censorship in 5 of the 6 cities, which resulted in the loss of approximately 1% of the reported person-years of follow-up; the reanalysis team updated the Six Cities cohort to include the missing person-years of observation, resulting in the addition of 928 person-years of observation and 14 deaths. The reanalysis team was able to reproduce virtually all of the original numerical results, including the 26% increase in all-cause mortality in the most polluted city (Stubenville, OH) as compared to the least polluted city (Portage, WI). The audit and validation of the Harvard Six Cities Study conducted by the reanalysis team generally confirmed the quality of the data and the numerical results reported by the original investigators. The discrepancies noted during the audit were not of epidemiologic importance, and did not substantively alter the original risk estimates associated with particulate air pollution, nor the main conclusions reached by the original investigators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".