[Commentary] A HARD NATURAL EXPERIMENT TO FATHOM: A TESTAMENT TO THE INCREASING DIFFICULTY OF CONDUCTING RELIABLE SURVEYS?
Notice bibliographique
Résumé
The paper by Mäkeläet al. [1] adds to the important tradition of studying ‘natural experiments’ in alcohol policy, a tradition that is especially strong in the Scandinavian countries [2]. Had the results followed the almost invariable pattern that reduced taxes lead to increased consumption, there would have been much less over which to discuss and puzzle. The authors are to be applauded for publishing these unexpectedly negative results. Arguably, studies such as this have unique value by forcing the field to question received wisdom and, I suggest, also traditional research methods. In discussing the results it is important to bear in mind the backdrop of a multitude of studies using many different designs, nearly all of which have indicated that alcohol behaves like other commodities, in so far as its consumption is responsive to price changes and hence manipulation of price is of major public health importance [3-6]. The main objective of this particular study was to study differential responsiveness to price changes among different subpopulations to confirm whether or not age, gender, income and the level of alcohol consumption influence substantially the degree of this responsiveness. However, the survey-based measures of drinking mainly failed to replicate the changes that were both expected from a reduction in taxation in Finland and Denmark (i.e. an increase in volume of consumption) and which were also observed in national sales data. The latter showed that a substantial reduction in Danish spirit taxes was associated with a 16% increase in spirits consumption (but 2% reduction in overall alcohol consumption as beer and wine sales decreased) and also that across-the-board reductions in Finnish liquor taxes resulted in a 10% increase in overall consumption. The panel design surveys used in Makela et al.'s study [1], however, found significant decreases in overall alcohol consumption in these two countries following the tax reductions, while the repeated cross-sectional surveys found no significant overall changes (although they had less power to detect these). A tempting conclusion at this point is that the panel survey design was not sufficiently sensitive to the population-level changes in consumption suggested by sales data, and hence questions of differential responsiveness by population subgroups could not be addressed. There are two broad areas of complexity which may have contributed to this discordance between survey and sales data. First, the natural experiment itself is not straightforward: (i) these are small neighbouring countries with a strong tradition of cross-border trade in alcohol (e.g. [7]); (ii) in Denmark only spirits taxes were reduced and this took place only three-quarters of the way through the first comparison year; (iii) changes to the cross-border allowances between Finland and Estonia were made in the middle of one of the second comparison years, 2004; (iv) there was a substantial relaxation of travellers' allowances introduced at the beginning of 2004 in Finland and the control country, Sweden. Secondly, there are a number of reasons for doubting the reliability and validity of the surveys conducted and the comparisons between them made to evaluate these changes: (i) as is increasingly common, response rates to the surveys, especially telephone surveys, were low (typically around 50%); (ii) research has shown repeatedly that data based on surveys drastically underestimate true consumption in populations, partly by under-sampling some of the heavier drinkers (e.g. [8]); (iii) the authors note the problem of regression to the mean in studies of changes in drinking patterns using a panel design, and while they attempted to adjust for these there may still have been differential effects operating across different age, gender and alcohol consumption groups; (iv) the authors note that differential dropout of subjects may occur in different subgroups of interest in panel studies which may be harder to assess precisely; (v) the results obtained from the panel and the cross-sectional surveys were mainly inconsistent with each other in relation to the size and direction of changes in consumption between 2003 and 2004, suggesting the presence of random errors and/or biases differentially effecting the two survey methods. In summary, I suggest that the sizes of actual changes in consumption that occurred, following what appeared to be dramatic changes in tax levels, may not have been substantial enough to enable fine-grained analyses of differential effects on populations subgroups, given the multiple difficulties of accurate measurement over time using population surveys. Future studies of natural experiments in tax policy changes may need to focus upon larger changes in consumption, perhaps over longer time-periods, with ‘cleaner’ comparisons possible between intervention and control countries, using cross-sectional surveys with larger sample sizes and sampling strategies which result in much higher response rates than 50%. Admittedly, this is setting the bar very high and at a standard that is increasingly hard to achieve in most countries, especially with telephone surveys (e.g. [9]). The use of measures of population rates of alcohol-related harm from morbidity and/or mortality data is also recommended (e.g. [10]). Overall, these mixed results underscore the value of employing multiple data sources to interpret changes in the drinking behaviour and the need for caution when interpreting observed changes over time assessed by population surveys alone. I am grateful to Scott Macdonald for helpful comments on a first draft of this commentary.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,007 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,002 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».