Bibliographic record
Abstract
Two papers in this issue demonstrate the utility of observational methods for evaluating the health effects of policies.1,2 These are helpful as they show the value of such approaches and also identify some of the methodological and other challenges. They are also timely because debates about the need for more evaluations of public health policies and other types of intervention research have moved on significantly in the past 10 years. Macintyre recently noted that commentaries from the 1990's had pointed to the lack of robust evidence to support social and public health policies in the UK.3 These highlighted the need for more robust and relevant evidence, and noted the lack of evaluations, particularly around health inequalities. This debate about gaps in the evidence base seems to have developed rapidly in recent years into discussions about the methodological implications of such gaps, and the challenges in producing new, reliable evidence. One important piece of methodological guidance to public health researchers in the UK emerged in 2000: this was the first edition of the Medical Research Council's (MRC) Guidance on complex interventions, which focused on the development and evaluation of complex public health interventions, and randomized controlled trials in particular.4 The second edition of this guidance, which appeared 8 years later, however also considered the place of other types of evaluative research, including the use of time series analyses for evaluations of the impact of natural experiments, with detailed examples.5 The MRC has followed this up even more recently by exploring the need for guidance on the evaluation of natural experiments, and a recent report of a workshop makes the point that in some circumstances an ideal study design will not be possible. Observational methods will therefore be necessary, though the findings may often need treated with caution.6 Although the biases in observational designs are well known, it is likely that in many cases the only evidence available will be this type of weaker evidence. Nonetheless, it plays an essential part. In public health, study design is frequently confounded with intervention type—so placing restrictions on the study designs that can be used in the interests of methodological rigour also inadvertently places a restriction on the types of intervention that can be evaluated; the best becomes the enemy of the good.7 For example, a systematic review which examined the effects of transport-related interventions to improve physical activity found that population-level interventions were less likely than individual-level interventions to have been studied using the most rigorous study designs.8 This review would have missed almost all population-level evidence if it had only included randomized controlled trials. The authors of these two studies on alcohol-related harms suggest caution in interpreting the findings, and Gustafsson points out the major limitations of such quasi-experimental designs, in particular that the control and intervention groups cannot be randomly selected.2 Donald Campbell, a major exponent of such methods, argued that the advocated strategy is therefore ‘ … not to throw up one's hands and refuse to use the evidence … but to generate as many plausible rival hypotheses, and then do the supplementary research … which would reflect on these rival hypotheses’.9 This is an important point and highlights the difference between studies which assess the effectiveness of narrowly defined, simpler interventions with respect to a few health outcomes, and observational evaluations of social policies which very often have a much broader goal—that is, to identify the range and nature of the many health and non-health impacts which follow the implementation of an intervention. Given the difficulty in drawing robust causal inferences from the latter type of study, the need to test alternative hypotheses (as Campbell suggests) is clear. It is also important to see the findings of such evaluations as far from definitive, but rather as only one smaller contribution to a larger, ever-cumulating evidence base. The results of evaluations of natural experiments are therefore rarely definitive, and only make sense when set alongside the wider population of similar studies which have examined similar interventions in different settings. To take a similar example, a single evaluation of the effects of smoking bans—which have often been evaluated as natural experiments—would not be persuasive. However, corroboration from numerous studies of similar designs gives greater confidence in the findings (though in this case it has also recently been suggested that there is significant variation between the studies, due to biases in estimates of the effects of the ban).10 This again is a reminder to take the findings of such studies in the context of other similar studies. It is probably helpful that debates about lack of evidence have developed into debates about the role of appropriate methods, and about the appropriate place of experimental and observational methods. Economist Julian Reiss has noted that in the case of natural experiments, economists divide into ‘sinners’ and ‘preachers’ on the basis of their theoretical and methodological preferences.11 There is a risk that public health researchers too become polarized into sinners and preachers—that is, trialists and non-trialists. In reality many evaluations of large-scale policies—like alcohol policies—require a certain amount of sinning, and this is appropriate if the public health evidence base is to develop (as long, of course, as we learn from our sins). It is highly likely that real advances in our knowledge about improving health and reducing health inequalities will come from such studies. Government policies are the major upstream social determinants of public health and it is unlikely that government departments will see large increases in the use of randomization in the near future. Observational evidence will therefore dominate, and there will never be a perfect public health base derived from robust trials. However, researchers can continue to support further development and can argue for the greater adoption of experimental methods, which Macintyre argues in public health are both more possible than many objectors think, and have a greater power to convince.2 In the meantime, evaluations of natural experiments have an essential role to play, not just in understanding impacts but also assessing impacts within different contexts, settings and population subgroups. Conflict of interest: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.060 | 0.266 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.005 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.006 | 0.008 |
| Scholarly communication | 0.007 | 0.010 |
| Open science | 0.009 | 0.004 |
| Research integrity | 0.074 | 0.064 |
| Insufficient payload (model declined to judge) | 0.013 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".