MétaCan
Menu
Retour à la cohorte
Enregistrement W1505229351 · doi:10.1111/j.1360-0443.2011.03375.x

Commentary on Brennan <i>et al</i>. (2011): Towards more interpretable evaluations

2011· letter· en· W1505229351 sur OpenAlexaff
Kathryn Marie Graham

Notice bibliographique

RevueAddiction · 2011
Typeletter
Langueen
DomaineMedicine
ThématiqueSubstance Abuse Treatment and Outcomes
Établissements canadiensCentre for Addiction and Mental Health
Organismes subventionnairesnon disponible
Mots-clésPsychological interventionComparabilitySAFERIntervention (counseling)PsychologyAggressionOutcome (game theory)Behavior changeInjury preventionApplied psychologyPoison controlMedicineClinical psychologySocial psychologyMedical emergencyPsychiatryComputer securityComputer science

Résumé

récupéré en direct d'OpenAlex

Brennan et al.'s review 1 highlights the need for more complex intervention models and multiple outcome measures in evaluations of interventions with licensed premises. In this commentary, I discuss applications of basic evaluation design that can address this need and improve comparability and meaningfulness of research in this area. Brennan et al. 1 reviewed evidence for two outcomes: disorder and severe intoxication. However, effectiveness of interventions may differ depending on the outcome being examined. For example, responsible beverage service (RBS) programs may reduce intoxication under some circumstances (see 2), but are unlikely to reduce violence significantly because the effects of RBS on intoxication tend to be small, and intoxication is only one among several factors contributing to bar violence 3, 4; other factors include staff behavior 5, environmental conditions 3, 4 and behavioral norms 6. Correspondingly, the Safer Bars program, which focuses exclusively on managing problem behavior (not RBS), reduced moderate–severe aggression 7, but would not necessarily be expected to reduce intoxication. Both interventions and evaluations in this area could benefit from using explicit logic models linking measures of implementation to mediating variables and ultimate outcomes. As an example, the Safer Bars evaluation found significant improvement in knowledge and attitudes among bar staff who received training 8, but no significant environmental changes 5, supporting the explanation that staff training, not environmental change, was the primary mechanism by which the reduction in violence was achieved. This explanation was supported further by 12-month follow-up data indicating that (a) venues that had lower turnover of Safer Bars-trained managers and security staff had lower rates of aggression than did venues with higher turnover 7, and (b) managers reported making few environmental changes following the program. Community mobilization projects, in particular, often include multiple interventions undertaken simultaneously, making it difficult to identify effective components. A model identifying mediating and moderating variables can contribute to a better understanding of how multiple-component programs work. In terms of moderating variables, for example, Wagenaar et al. 9 hypothesized that their community intervention would have a greater effect on 18–20-year-olds (who were the primary target) than on 15-17-year-olds; however, they found similar effects for both age groups, suggesting that the intervention may have had an effect on community norms as well as an effect attributable to the development of age-specific policies. The logic model can also identify unrealistic outcome expectations from interventions limited in scope or duration. For example, multi-component large-scale community interventions may have substantial and measurable community-level impacts, whereas bar-level interventions are unlikely to have effects that can be measured at the community level, even if significant at the bar level, unless delivered community-wide. The STAD project (Stockholm Prevents Alcohol and Drug Problems) illustrated the importance of intervention potency and duration with a 10-year time-frame, during which the effects of this multi-component program actually increased over time 10, in contrast to most other interventions where effects have tended to diminish over time 5. Police arrest data are often used to estimate an intervention's effect on crime; however, these data are unlikely to provide a valid assessment when the intervention involves changes in police activity. For example, a police-led 11 and a community intervention 12 involving enhanced policing both found an increase in alcohol-related arrests following the intervention; however, the community study 12 also found a significant decrease in emergency room visits for assault, suggesting a real reduction in assaults mediated possibly by enhanced policing. Evaluations in this area sometimes ignore well-known threats to internal and external validity 13. For example, when an intervention has occurred in response to a natural spike in problems, improvement following the intervention could be attributable to regression to the mean, a design problem not solved by the use of a non-equivalent comparison area. This explanation could account for the dramatic decline in assault rates following the first year of the Geelong Accord 14, 15, although data from subsequent years suggest a possible real effect of the Accord. To conclude, there are many methodological challenges to implementing and evaluating interventions in the community. As Brennan et al. suggest, one solution to understanding results from a diverse set of studies is research to ‘unify disparate measures of harm so that studies can be compared’. However, more interpretable findings are possible if interventions and evaluations are based on logic models that clearly identify the process by which the intervention is expected to work and distinguish between measures of implementation, mediation and outcomes. In addition, reviews could also use a logic model to frame their assessment of existing findings. Such an approach to reviews would allow for meaningful reviews to be conducted without unnecessarily restrictive criteria that require the exclusion of important studies (e.g. 16, 17) because they do not meet specific design criteria. None.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,075
score de la tête « metaresearch » (Gemma)0,310
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,079
Score d'incertitude au seuil0,396

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0750,310
Méta-épidémiologie (sens strict)0,0030,003
Méta-épidémiologie (sens large)0,0040,005
Bibliométrie0,0040,004
Études des sciences et des technologies0,0090,013
Communication savante0,0110,014
Science ouverte0,0150,007
Intégrité de la recherche0,0790,087
Charge utile insuffisante (le modèle a refusé de juger)0,0110,011

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,034
Tête enseignante GPT0,314
Écart entre enseignants0,280 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2011
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAddictionMême sujetSubstance Abuse Treatment and OutcomesTravaux en français237 207