MétaCan
Menu
Retour à la cohorte
Enregistrement W1574901265 · doi:10.1111/j.1360-0443.2007.02130.x

[Commentary] FIDDLING WHILE ROME BURNS? BALANCING RIGOUR WITH THE NEED FOR PRACTICAL KNOWLEDGE

2008· article· en· W1574901265 sur OpenAlexaff
Kathryn Graham

Notice bibliographique

RevueAddiction · 2008
Typearticle
Langueen
DomaineHealth Professions
ThématiqueCommunity Health and Development
Établissements canadiensCentre for Addiction and Mental Health
Organismes subventionnairesnon disponible
Mots-clésRigourNothingVariety (cybernetics)ConstructiveStandardizationConsistency (knowledge bases)Public relationsIntervention (counseling)PsychologySimplicityEngineering ethicsManagement scienceMedicinePolitical scienceComputer scienceEngineeringNursingLawEpistemology

Résumé

récupéré en direct d'OpenAlex

Around the world, communities are searching for and adopting strategies and programmes to address problems related to drinking in licensed premises [1-4]. Most of the programmes and strategies that are adopted are essentially unevaluated; or if they have an evaluation component, it is often of the ‘Let’s put on a show in the barn' variety. That is, the evaluation is conducted (often on a shoestring budget) by people with no basic knowledge of how to design an evaluation in a way that will allow them to draw meaningful or valid conclusions about the programme's impact. Nevertheless, many local communities are prepared to adopt unproven approaches rather than do nothing at all. The research community, on the other hand, demands the highest level of methodological rigour; consequently, studies are left with nothing to offer if promising approaches do not meet rigorous evaluation criteria. For example, the study by Toomey et al. (this issue [5]) included a thoughtful, well-developed, research-based intervention [6] with an excellently designed and implemented evaluation conducted by highly trained, skilled and experienced researchers. However, with null findings on programme effectiveness, the study could not offer any constructive recommendations for communities looking for solutions to addressing problems related to over-serving in licensed premises. This is not intended to be a critique of their exemplary research, as their standards and methods meet the universally endorsed gold standards for programme evaluation described by Cook & Campbell [7]. Rather, their study raises the question of whether we need to take into consideration practical needs, as well as rigour when we design evaluations of community interventions. So how can research increase practical knowledge while still maintaining high standards of scientific rigour? The first step might be to take a more nuanced approach to conceptualizing outcomes. For example, it is possible that a significant effect of the intervention (relative to the comparison group) might have been found had the study employed a more sensitive outcome measure. The dichotomous outcome measure (served/not served) employed by Toomey et al. [5] based on one visit by a pseudo-patron feigning intoxication would be insensitive to minor improvements in reducing service to intoxicated people. For example, some of the non-refusal strategies that Toomey et al. refer to as ‘mild interventions, such as offering alcohol-free beverages’[8-12] may well have the impact of reducing overall intoxication levels of patrons and related harms, despite no measurable impact on service refusal. In addition, adopting a broader view of outcomes such as measuring problems associated with intoxication [13] (e.g. violence, vandalism and disorder, driving, injury) as well as serving practices might provide additional valuable information about the effects of policy-oriented approaches such as alcohol risk management (ARM). Secondly, assessment of statistical significance may need to be reconsidered in the context of evaluations of community prevention initiatives. The problem with the ‘rigour or nothing’ approach is most apparent when an effect is found that is not big enough to meet the criterion for statistical significance. By focusing on avoiding Type 1 error, we increase the odds of committing Type 2 error, with the result that potentially useful approaches to prevention may never become known to those in the community who are seeking solutions. Meta-analyses and Cochrane reviews are useful tools for addressing the problem of interpreting significant findings (or lack thereof) from a single study, especially in fields of research where it is possible to conduct numerous trials relatively cheaply. However, there are fewer opportunities for cross-study analyses of evaluations of community interventions, which are relatively rare due to the long time-frame required and high costs of implementation. Perhaps we should tolerate a higher Type 1 error (e.g. being wrong one in 10 times rather than one in 20) when we conclude that a programme is effective, as long as there is clear evidence that there is no negative impact of the programme. Perhaps the solution for evaluation of preventive interventions is to adopt an ordinal rather than a dichotomous method for interpreting significance when results are in the predicted direction—for example: P < 0.05 could count as ‘proven effective’ according to usual norms; P < 0.10 and ≥ 0.05 could be counted as ‘evidence of probable effectiveness; and P < 0.20 and ≥ 0.10 might be classified as ‘promising and no evidence of negative effects’. Although the lack of impact on alcohol policies supports the conclusions of Toomey et al.[5] that the ARM programme did not have the desired effect, it is interesting that there are now two evaluations of ARM [5, 14] which found immediate increases in refusal of service, although in both cases the effect did not meet the criterion of P < 0.05 for statistical significance. Taken together, the two studies provide evidence of a small but consistent effect. This is not to suggest that such training would be sufficient to eliminate all serving to intoxication, but perhaps its potential usefulness should also not be disregarded. Thirdly, qualitative and descriptive research needs to be included in evaluations both to improve interventions as well as to understand more clearly why a programme failed to have an impact. For example, findings from studies that have examined why staff and management have difficulties refusing service to intoxicated patrons [15] could be incorporated into the development of more effective approaches to preventing intoxication and more nuanced outcome measures. It would have been interesting in the Toomey et al. study [5] to have asked bar managers why their staff were still serving intoxicated patrons, despite participation in the training. One also wonders why adoption of policies by experimental premises was no better than the adoption rate for control premises, given that this was the focus of the intervention. Traci Toomey and her colleagues have conducted considerable work [5, 14] describing explanatory and process factors, and perhaps have data addressing these issues that they will be publishing elsewhere. Finally, the relevance of scientific research generally to community applications might be improved by expanding the study time-frame, including implementing long-term interventions. For example, the most successful responsible beverage service project to date, the STAD (Stockholm Prevents Alcohol and Drug Problems) project in Sweden, had a 10-year time-frame, and the impact of the programme showed a steady increase over time [16, 17]. This means that we need to change the current research context in many countries, especially the time-line of grant funding, that typically does not permit this kind of long-term implementation and evaluation. This commentary is not intended as a criticism of this well-conducted research. Rather, I have used this paper to raise the general issue of increasing the potential for applicability of research findings while maintaining high methodological standards. With communities and the general public increasingly recognizing the need for evidence-based interventions, perhaps it is time for researchers to change how we design community prevention research so that we can increase practical relevance of findings even when we are not able to demonstrate a statistically significant impact on the primary outcome measure.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesÉtudes des sciences et des technologies
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,236
Score d'incertitude au seuil0,998

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0030,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,096
Tête enseignante GPT0,408
Écart entre enseignants0,312 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations9
Publié2008
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAddictionMême sujetCommunity Health and DevelopmentTravaux en français237 207