Simulation study on the validity of the average risk approach in estimating population attributable fractions for continuous exposures
Bibliographic record
Abstract
BACKGROUND: The population attributable fraction (PAF) is an important metric for estimating disease burden associated with causal risk factors. In an International Agency for Research on Cancer working group report, an approach was introduced to estimate the PAF using the average of a continuous exposure and the incremental relative risk (RR) per unit. This 'average risk' approach has been subsequently applied in several studies conducted worldwide. However, no investigation of the validity of this method has been done. OBJECTIVE: To examine the validity and the potential magnitude of bias of the average risk approach. METHODS: We established analytically that the direction of the bias is determined by the shape of the RR function. We then used simulation models based on a variety of risk exposure distributions and a range of RR per unit. We estimated the unbiased PAF from integrating the exposure distribution and RR, and the PAF using the average risk approach. We examined the absolute and relative bias as the direct and relative difference in PAF estimated from the two approaches. We also examined the bias of the average risk approach using real-world data from the Canadian Population Attributable Risk of Cancer study. RESULTS: The average risk approach involves bias, which is underestimation or overestimation with a convex or concave RR function (a risk profile that increases more/less rapidly at higher levels of exposure). The magnitude of the bias is affected by the exposure distribution as well as the value of RR. This approach is approximately valid when the RR per unit is small or the RR function is approximately linear. The absolute and relative bias can both be large when RR is not small and the exposure distribution is skewed. CONCLUSIONS: We recommend that caution be taken when using the average risk approach to estimate PAF.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.082 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".