MétaCan
Menu
Back to cohort

[Commentary] A HARD NATURAL EXPERIMENT TO FATHOM: A TESTAMENT TO THE INCREASING DIFFICULTY OF CONDUCTING RELIABLE SURVEYS?

2008· letter· en· W1525165108 on OpenAlexaff
Tim Stockwell

Bibliographic record

VenueAddiction · 2008
Typeletter
Languageen
FieldHealth Professions
TopicHealthcare cost, quality, practices
Canadian institutionsUniversity of Victoria
Fundersnot available
KeywordsConsumption (sociology)Value (mathematics)Natural experimentEconomicsPublic economicsMultitudeAlcohol consumptionDemographic economicsPsychologyPositive economicsMedicineSociologyPolitical scienceLawAlcoholSocial science

Abstract

fetched live from OpenAlex

The paper by Mäkeläet al. [1] adds to the important tradition of studying ‘natural experiments’ in alcohol policy, a tradition that is especially strong in the Scandinavian countries [2]. Had the results followed the almost invariable pattern that reduced taxes lead to increased consumption, there would have been much less over which to discuss and puzzle. The authors are to be applauded for publishing these unexpectedly negative results. Arguably, studies such as this have unique value by forcing the field to question received wisdom and, I suggest, also traditional research methods. In discussing the results it is important to bear in mind the backdrop of a multitude of studies using many different designs, nearly all of which have indicated that alcohol behaves like other commodities, in so far as its consumption is responsive to price changes and hence manipulation of price is of major public health importance [3-6]. The main objective of this particular study was to study differential responsiveness to price changes among different subpopulations to confirm whether or not age, gender, income and the level of alcohol consumption influence substantially the degree of this responsiveness. However, the survey-based measures of drinking mainly failed to replicate the changes that were both expected from a reduction in taxation in Finland and Denmark (i.e. an increase in volume of consumption) and which were also observed in national sales data. The latter showed that a substantial reduction in Danish spirit taxes was associated with a 16% increase in spirits consumption (but 2% reduction in overall alcohol consumption as beer and wine sales decreased) and also that across-the-board reductions in Finnish liquor taxes resulted in a 10% increase in overall consumption. The panel design surveys used in Makela et al.'s study [1], however, found significant decreases in overall alcohol consumption in these two countries following the tax reductions, while the repeated cross-sectional surveys found no significant overall changes (although they had less power to detect these). A tempting conclusion at this point is that the panel survey design was not sufficiently sensitive to the population-level changes in consumption suggested by sales data, and hence questions of differential responsiveness by population subgroups could not be addressed. There are two broad areas of complexity which may have contributed to this discordance between survey and sales data. First, the natural experiment itself is not straightforward: (i) these are small neighbouring countries with a strong tradition of cross-border trade in alcohol (e.g. [7]); (ii) in Denmark only spirits taxes were reduced and this took place only three-quarters of the way through the first comparison year; (iii) changes to the cross-border allowances between Finland and Estonia were made in the middle of one of the second comparison years, 2004; (iv) there was a substantial relaxation of travellers' allowances introduced at the beginning of 2004 in Finland and the control country, Sweden. Secondly, there are a number of reasons for doubting the reliability and validity of the surveys conducted and the comparisons between them made to evaluate these changes: (i) as is increasingly common, response rates to the surveys, especially telephone surveys, were low (typically around 50%); (ii) research has shown repeatedly that data based on surveys drastically underestimate true consumption in populations, partly by under-sampling some of the heavier drinkers (e.g. [8]); (iii) the authors note the problem of regression to the mean in studies of changes in drinking patterns using a panel design, and while they attempted to adjust for these there may still have been differential effects operating across different age, gender and alcohol consumption groups; (iv) the authors note that differential dropout of subjects may occur in different subgroups of interest in panel studies which may be harder to assess precisely; (v) the results obtained from the panel and the cross-sectional surveys were mainly inconsistent with each other in relation to the size and direction of changes in consumption between 2003 and 2004, suggesting the presence of random errors and/or biases differentially effecting the two survey methods. In summary, I suggest that the sizes of actual changes in consumption that occurred, following what appeared to be dramatic changes in tax levels, may not have been substantial enough to enable fine-grained analyses of differential effects on populations subgroups, given the multiple difficulties of accurate measurement over time using population surveys. Future studies of natural experiments in tax policy changes may need to focus upon larger changes in consumption, perhaps over longer time-periods, with ‘cleaner’ comparisons possible between intervention and control countries, using cross-sectional surveys with larger sample sizes and sampling strategies which result in much higher response rates than 50%. Admittedly, this is setting the bar very high and at a standard that is increasingly hard to achieve in most countries, especially with telephone surveys (e.g. [9]). The use of measures of population rates of alcohol-related harm from morbidity and/or mortality data is also recommended (e.g. [10]). Overall, these mixed results underscore the value of employing multiple data sources to interpret changes in the drinking behaviour and the need for caution when interpreting observed changes over time assessed by population surveys alone. I am grateful to Scott Macdonald for helpful comments on a first draft of this commentary.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Research integrity
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.173
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0070.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0020.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0010.004
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.526
GPT teacher head0.499
Teacher spread0.027 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueAddictionSame topicHealthcare cost, quality, practicesFrench-language works237,207