MétaCan
Menu
Back to cohort
Record W4417020642 · doi:10.1111/add.70273

HEALing communities study results, questions and implications

2025· editorial· en· W4417020642 on OpenAlexaboutno aff
Jonathan P. Caulkins

Bibliographic record

VenueAddiction · 2025
Typeeditorial
Languageen
FieldMedicine
TopicOpioid Use Disorder Treatment
Canadian institutionsnot available
FundersNational Science Foundation
KeywordsOpioid overdose(+)-NaloxoneOpioid use disorderHarm reductionHeroinOpioidDrug overdoseAddictionPolysubstance dependence

Abstract

fetched live from OpenAlex

The HEALing Communities Study (HCS) was a $350 million 4-year multi-site, community-level, cluster-randomized wait-list controlled trial of evidence-based practices for reducing opioid overdose deaths. It did not produce statistically significant reductions in deaths (its central outcome), treatment uptake or behavioral health service delivery, but it did reduce stigma. This disappointing result should prompt serious reflection within our community. The $350 million HCS was the ‘largest implementation science study ever funded in addiction research’ [1]. The evaluation was rigorous, with 67 communities in four states randomly assigned to the Communities that HEAL (CTH) treatment or to the control group. The stated goal was ‘to reduce opioid overdose deaths by 40% in three years’ predicated on a belief that ‘opioid overdose deaths are largely preventable’ [1]. HCS embodied the field's best wisdom by including (1) community engagement; (2) communication campaigns to increase awareness and demand for evidence based practices (EBPs) and to reduce stigma against people with opioid use disorder (OUD) and against medications for treating opioid use disorder (MOUD); and (3) a requirement that communities implement EBP for (3a) overdose education and naloxone distribution (OEND), (3b) MOUD and (3c) safer prescribing of opioid analgesics that could ‘significantly reduce opioid overdose deaths in a relatively short period of time’ [1]. Regarding the central outcome, there was a statistically not-significant 8% reduction in the opioid overdose death rate (P = 0.30) [2] and the overall overdose death rate (P = 0.26) [3]. Secondary outcomes included a statistically significant 37% decline in deaths from opioids combined with a psychostimulant other than cocaine, and small and not statistically significant reductions in deaths from opioids plus cocaine (6%) and opioids with benzodiazepine (1%). Outcomes for additional aims (e.g. testing the study's conceptually driven framework) are less easily summarized. Related studies found no statistically significant effect on (1) initiation, retention, and linkage to MOUD; (2) the rate of waivered practitioners or active prescribing of buprenorphine; or (3) the rate of individuals receiving behavioral health services reflected in Medicaid claims [4-6]. However, ‘the CTH intervention significantly changed stakeholders' perceived community stigma toward OUD and MOUD’ (P = 0.0007 and P = 0.0066, respectively) [7]. These results challenge confidence that CTH's recipe of community engagement; reducing stigma; and EBP for MOUD, OEND and safer prescribing necessarily produce major changes in death and other health outcomes. Researchers, policy makers and others who had that confidence need to adjust their beliefs or identify reasons why the trial failed. Three main conjectures have been offered for why CTH could have failed even if its approach remains sound. First, the intervention began shortly before coronavirus disease (COVID), hampering deployment. However, 235 EBPs were implemented by the start of the evaluation's comparison year [2], which extended through 30 June 2022. Additionally, implementation proceeded enough to reduce stigma [7], although COVID could have interfered more with healthcare delivery than with stigma reduction efforts. Second, ‘change in the illicit drug market may have reduced the effectiveness of the intervention, because fentanyl became a more prevalent opioid … [and] we do not know whether surges in fentanyl use over time were similar across communities’ [2]. Figure 1 plots the proportion of illegal opioid observations that were fentanyl for counties containing control communities, counties containing intervention communities and other counties in the intervention states (see Caulkins and Giri for more on the data) [8]. Fentanyl's spread was similar in all three. If anything, the spread was greater between baseline and evaluation for control counties, which could have enhanced not attenuated the apparent effectiveness of CTH. Therefore, although CTH may be less effective in the fentanyl era than in previous times, it is not clear that the particulars of fentanyl's spread produced a false negative result in HCS's evaluation of CTH. Third, ‘the HCS timeline and reach of selected EBPs may have been insufficient’ [4]; that is, EBPs may produce change—but not quickly—and/or the dose was too small. Barocas et al.'s [9] analysis of the economic value of resources deployed for interventions may support the latter idea. Although the HCS study budget was nearly $350 million, payments made by HCS to implement CTH totaled only $37.5 million. In-kind resources (e.g. community members' time) and non-HCS financial support added another $26.4 million. Therefore, the total economic value of resources devoted to CTH was $63.8 million or just $1.93 million per intervention site. Further, half of those resources were devoted to community engagement, and more to communications campaigns, leaving only one-third ($668 000 per site) for implementing EBP strategies. In theory, a fourth possibility is that because CTH empowered each community to make their own selections from the EBP list, the communities could have selected unwisely. If they selected options on the weaker end of the evidence-based list, that could have diluted the total impact. For those inclined to adjust beliefs about CTH in light of HCS results, there are many possibilities. To mention four possibilities that I ponder: (1) Perhaps the power of stigma-reduction efforts has been over-sold, since HCS reduced stigma but did not produce the other outcomes. (Note: RCTs evaluating stigma reduction's effects on distal outcomes, like death, are scant compared to RCTs on MOUD.) (2) Maybe CTH has high efficacy in small studies, but does not scale-up well. (3) Perhaps EBP are cost-effective (with only $668 000 invested per site in programming, saving a single life per site would cost-justify them), but do not save enough lives to bend the curve appreciably at a population level. (4) Perhaps EBP could bend the curve, but only with much greater investments than planners designing the HCS presumed it would take. Given the HCS results, I also wonder what did produce the much greater than 8% reduction in deaths that occurred between mid-2023 and mid-2025 for the United States as a whole and also for Canada. I do not advocate one view over another or wish for readers necessarily to agree with me. I do, though, suggest that we collectively engage with this challenge and be open to some change in thinking. If we do not change our minds in some way, can we ask taxpayers for another $350 million study—when that sum could instead fund CTH-scale EBP programming in approximately 500 communities? The public health community often urges people to heed scientific evidence on matters ranging from vaccines to dietary advice to COVID mitigation. Will we heed the HCS study evidence enough to change our minds in some way? Jonathan P. Caulkins: Writing—original draft (lead); writing—review and editing (lead). The author thanks Keith Humphreys, Beau Kilmer, Peter Reuter and two anonymous referees for helpful comments on earlier drafts. None. The data that support Figure 1 were made available by the HIDTA PMP program. They cannot be shared directly by the author but should be available from HIDTA PMP to anyone wishing to replicate the figure.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: Editorial
Teacher disagreement score0.182
Threshold uncertainty score0.787

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.016
GPT teacher head0.328
Teacher spread0.312 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueAddictionSame topicOpioid Use Disorder TreatmentFrench-language works237,207