Testing the Harvesting Hypothesis by Time-Domain Regression Analysis. I: Baseline Analysis
Bibliographic record
Abstract
Although the association between air pollution and daily mortality is well established, the mechanisms by which air pollution results in excess mortality are not yet well understood. In particular, there exists debate over whether air pollution has a direct effect on mortality in the general population or simply shortens the life span of frail individuals, a hypothesis referred to as "harvesting." The goal of this investigation is to test the harvesting hypothesis using the time-domain regression method of Dominici et al. (2003a). We conducted simulations based on a two-compartment model that divides the population into a larger group of healthy individuals and a frail subpopulation. Death from air pollution is assumed to take place in two steps, by first moving from healthy population to the frail pool, then death with probability related to the level of air pollution. Using time-domain analysis, we seek to identify data patterns that would be characteristic of harvesting under different scenarios. For a pure harvesting model, time-domain analysis indicates that mortality is associated with a short-term air pollution episode of less than 2 d if the mean residency time in the frail pool is short. If both entrants and deaths depend on the level of air pollution and the rates of entry to and exit from the frail pool are about the same, the log relative risk estimates are essentially unchanged at all time scales. If pollution affects mortality in the frail pool more than entrants, larger effects will occur at shorter time scales.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.048 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.004 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".