Notice bibliographique
Résumé
Global warming, health inequities, infectious disease pandemics, obesity: many of the world's most important health problems are complex, as are interventions proposed to attenuate their harmful effects. With this in mind, Smith and Petticrew's call for broader evaluations of public health interventions is welcomed.1 However, seeing the need for ‘macro-evaluations’ is the easy task. The hard task, as the authors acknowledge, is actually performing such studies. Despite a decade of similar advocacy, few macro-evaluations have been performed.2 Smith and Petticrew recommend that public health consider new methods, embracing collaboration with other disciplines, because the traditional micro-approach of public health is too narrow for the task. While we agree that public health needs to broaden its toolkit, we suggest there is much to learn from successful macro-evaluations studies already performed. We review two of our favourite studies and identify potential lessons related to the strength of evidence; scope of evaluation (extent of ‘macro-ness’ as defined by Smith and Petticrew); and dependence on leadership and stakeholder engagement. We hope others will discuss lessons learned from other studies. Whether the field is public health or other disciplines, most evaluations fall into one of the following categories: (1) experimental, where individuals or populations are randomly assigned by the investigators to receive the intervention; (2) observational, where the intervention is not determined by the investigators; and (3) modeled, where investigators simulate the introduction of an intervention or a combination of interventions, and various inputs can be manipulated to predict and examine a range of potential outcomes. Our first favourite is an experimental study of a Mexican incentive-based welfare program called ‘Oportunidades’ that provided investments in nutrition, health and education for young children living in low-income families.3 The program (or intervention) consisted of micronutrient-fortified food for women and children, cash transfers to families that were conditional on attendance at school, health care appointments, and a mandatory nutrition and health education session. The study by Rivera et al.3 described the nutritional impact in a subgroup of 347 communities that were randomized to the intervention immediately or after a 1-year delay. We learned two lessons and noted one drawback from this study. First, that it is possible to incorporate a high-quality intervention trial into a multi-component social program that is delivered at a massive scale;4 by 2004, the Oportunidades program covered 4.5 million families. Second, leadership from the highest levels of government is prerequisite for implementing innovative health policy that spans multiple ministries of government as well as the private sector.5,6 Therefore, we postulate that support from high-level leadership was instrumental in the Oportunidades intervention study. The drawback is that the study is still a ‘micro-evaluation’ from Smith and Petticrew's perspective. The intervention was essentially a single cause-effect and the main outcomes were biomedical markers of health (children's height and anaemia). Our second favourite is a modeling study by Woodcock et al.7 that assessed urban transportation and the environment. This study estimated the health effects of alternative urban land transport scenarios—lower-carbon-emission vehicles versus increased active travel versus a combination of the two. The authors examined the impact of these hypothetical policies on physical activity, air pollution and the risk of road traffic injury. Although only health outcomes were reported, the study included provisions to evaluate non-health outcomes such as economic growth. The lesson here is that modeling studies are ideally suited for macro-evaluation. Woodcock et al.'s7 study has most of the characteristics of a macro-evaluation, such as multiple sectors, disciplines and causal pathways. Modeling studies often require the involvement of ‘untraditional bedfellows’ that Smith and Petticrew encourage us to collaborate with, and are the cornerstone of many different disciplines' effort to describe the natural history and likely outcome of events, notably fields such as ecology, environmental sciences, engineering and economics. Macro-evaluative studies almost by definition require a wide range of study types and data, and it is helpful to look for ways to use existing studies and data to support this complex work. Modeling studies, such as the example of urban transportation, need data encompassing multiple viewpoints to describe population risk exposure, hazards or transitions from different health states, population and economic outcomes, physical and social structures and interactions, and so forth. However, modeling studies are only as robust as the evidence that goes into building the models; therefore, it is critical that models incorporate evidence from experimental and high-quality observational studies of individual policies or interventions. Thus, modeling studies combine individual micro-evaluations into a macro-evaluation. In this way, our most important lesson becomes apparent when you examine these studies together. Experimental studies generally provide the strongest evidence, and are important for building a case for policy effectiveness,2 but they tend to be micro- rather than macro-evaluations and are most dependent on leadership and stakeholder engagement. In contrast, modeling studies are often viewed as providing the weakest evidence but have the greatest potential to be macro-evaluations. Observational studies, in most cases, lie somewhere between the two extremes on these three dimensions of evidence, leadership and ‘macro-ness.’ All three study types are up for the task of macro-evaluation. More often, we should look at the wood, but a healthy wood is made from sound trees. Dr. Manuel holds a Chair in Applied Public Health from The Canadian Institute for Health Research and the Public Health Agency of Canada. Dr. Kwong is supported by a Career Scientist Award from the Ontario Ministry of Health and Long-Term Care and a Research Scholar Award from the Department of Family and Community Medicine, University of Toronto. The opinions, results and conclusions are those of the authors, and no endorsement by funding agents or the Ontario Ministry of Health and Long-Term Care or by the Institute for Clinical Evaluative Sciences is intended or should be inferred.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,021 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,007 | 0,005 |
| Communication savante | 0,004 | 0,005 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,056 | 0,065 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,011 | 0,007 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».