MétaCan
Menu
Back to cohort
Record W4387081864 · doi:10.1097/ede.0000000000001663

Empirical Challenges in Defining Treatments and Time in the Evaluation of Gun Laws

2023· article· en· W4387081864 on OpenAlexaffabout
Sam Harper, Arijit Nandi

Bibliographic record

VenueEpidemiology · 2023
Typearticle
Languageen
FieldSocial Sciences
TopicGun Ownership and Violence Research
Canadian institutionsMcGill University
Fundersnot available
KeywordsLawPolitical science

Abstract

fetched live from OpenAlex

Rigorous policy analysis is hard and fraught with many empirical challenges, which are often compounded for issues that are politically polarizing, such as gun policy in the United States. In this issue, Sharkey and Kang1 analyze over two decades worth of changes in US state-level gun policies and argue that they have been extremely successful in reducing gun deaths. Sharkey and Kang’s1 analysis has a number of commendable dimensions. They motivate their analysis with an important and timely policy question, which is whether changes in gun policies contributed to the decline in gun-related deaths during the period from 1991 to 2016. Although their primary focus is on estimates generated using a pooled outcome (all gun deaths) and a cumulative index of all state changes to gun laws, they helpfully provide separate estimates for several different classes of gun laws and examine their impact across different subsets of gun-related and nongun-related outcomes. They also conduct several sensitivity analyses (using an alternative database of laws, fixed effects, and instrumental variable analyses) to bolster the credibility of their causal claims, as well as providing the source code that was used to generate their figures and tables. These positive aspects notwithstanding, the preferred model specification and analytic framework that Sharkey and Kang1 use to generate their evidence also raise some important methodologic questions. WHAT IS THE TREATMENT? The large range of potential policies to regulate access to and use of guns creates serious challenges for evaluation.2 As noted above, Sharkey and Kang approach the measurement of gun laws through the simple lens of a cumulative index of changes in state policies. They code the adoption of each policy as either restrictive (+1) or permissive (−1), giving each state a cumulative value of the net changes from 1991 to 2016. As for any other causal question, to interpret Sharkey and Kang’s estimates of the policy index’s effects on deaths as causal, one must credibly assume consistency, exchangeability, and positivity,3 among other specification challenges discussed below. Sharkey and Kang’s coding scheme of the gun law index is transparent and simple, but it leads to ambiguous and difficult interpretations of the resulting regression estimates. First, amalgamating all state policies into an index forces the authors to make the strong assumption that any impacts are linear and additive, an assumption that could be potentially tested or relaxed with alternative coding or analytic methods. Second, and more importantly, the index leads to ambiguous interpretations because of plausible consistency violations. The consistency assumption is a way of linking the potential outcomes necessary for causal contrasts under different treatment scenarios to the observed outcomes.4 In Sharkey and Kang’s context, this means that for states with a given value a of their index, their gun death rate is identical to what it would be if we were to intervene to set their index to a. For example, in their Table, both Michigan and Mississippi have average 2012–2016 index values of +3, yet they arrived at this net index value through passing different sets of policies at different times. Consistency implies that there are not multiple versions of the treatment that would lead the potential outcomes to be different under alternative versions of the treatment. Critically for the authors, the consistency assumption here means that the effect of their indexed exposure would not differ based on which (net) policies states implemented to arrive at a given value of the index. Based on the existing and quite heterogeneous literature on specific gun laws and gun deaths in different populations,2 this seems like a near-heroic assumption. It also seems particularly unfortunate in this context since, unlike many other social exposures in epidemiology,5 making a credible case for the consistency assumption can be more straightforward for specific social policies implemented via legislation. However, even specific gun laws are challenging in this regard, since for the same “class” of laws (e.g., background checks or concealed-carry laws), there may be important state differences that create multiple versions of this specific treatment (e.g., background checks for dealers vs. purchasers, and for purchasers with criminal history vs. mental health problems, or “shall-issue” vs. “permitless carry” concealed-carry laws). This important challenge of multiple versions of treatment is compounded when entire classes of gun regulations are then pooled and treated as equivalent. Not only does the use of an index create potential problems for interpreting Sharkey and Kang’s estimated effects, but it also seems unlikely that using a pooled index would be of value to policymakers or other stakeholders in the debate over gun legislation. Much of the existing disagreement among policy experts centers around which specific policies are likely to be effective, and for which populations.6 The authors’ decision not only to pool all gun laws into an index but also to pool all outcomes irrespective of age, gender, or race may also reduce the utility of their results. Again, clearly, there are benefits to simplifying the analysis by looking at all laws and pooled outcomes, but this also comes at the cost of ignoring potential heterogeneity—particularly by race, gender, and age—that is valuable for understanding how gun laws may (or may not) work. For example, a number of studies have found that child access protection laws specifically reduce firearm suicides among youth,2,7 but this analysis shows little impact on the whole population, perhaps due to pooling all deaths. Similarly, past work has demonstrated substantial race and gender differences in the homicide decline,8,9 so ignoring the potential for differential policy impacts among population subgroups seems like a missed opportunity. WHEN IS THE TREATMENT? Leaving aside the challenges of treatment coding, the rationale for their preferred analytic strategy raises additional questions. Sharkey and Kang’s data include a 26-year time series of information on changes to state firearm regulations, firearm-related deaths, and covariates for 50 states, representing 1300 rows of data. Rather than analyzing the entire time series, they opt for a “change model,” in which the change in average firearm deaths in each state from 1991–1995 to 2012–2016 is modeled as a function of the net change in firearm regulations in that state over the same period. This reduces the number of datapoints in the primary analysis from 1300 to 50 rows. An alternative approach for evaluating the impact of changes in firearm regulations would have been to model the impact of year-to-year changes in firearm regulations on corresponding changes in firearm deaths within a state, which represents the difference-in-differences or “state and year fixed effects specification” that the authors considered in their sensitivity analyses. In linear specifications, these models estimate a variance-weighted average treatment effect on the treated, averaged over the entire postperiod.10 Sharkey and Kang’s rationale for the choice of the change model over the fixed effects specification is that the fixed effects model “requires speculative assumptions about how long a lag there is between the implementation of a state regulation and the effects on gun deaths,”1 whereas the change model considers the cumulative effect over the study period. Both approaches compare each state to itself, and in doing so account for time-fixed unobserved differences across states that may be associated with the adoption of firearm regulations and changes in firearm deaths. However, the two approaches differ in how they handle time, and there may be advantages to the fixed effects model compared with the change model that should be considered when designing impact evaluation studies. First, the fixed effects model, by specifying lags for the policy indicators, can be useful for establishing temporality between changes in firearm regulations and deaths. The change model, on the other hand, is agnostic about whether changes in firearm regulations preceded or followed changes in firearm deaths within a state and simply measures how strongly these changes are related. This limitation extends to the instrumental variable analyses that, by collapsing the annual data, presume that changes in firearm regulations are influenced by changes in the state’s contribution to total gun manufacturing. However, the reverse temporal ordering is also plausible, as the authors acknowledge in the eAppendix, so it is hard to say whether this instrument makes the case for causality any stronger. Second, although the fixed effects specification requires assumptions about the potential lag between a change in firearm regulations and subsequent changes in firearm deaths, it provides a framework for evaluating and modeling the dynamic effects of firearm reforms.11 This includes, for example, the ability to assess how soon after a change in gun laws we might observe changes in mortality–assessment of these “lagged” effects respond to policy-relevant questions6 about whether different types of regulations have immediate, delayed, temporary, and/or sustained effects. This seems important given that different law classes were implemented at different times. For example, nearly all of the “Castle Doctrine” laws were implemented since the mid-2000s, after the steep decline in gun deaths, but many of the “minimum age” or background check laws were implemented much earlier. With the fixed effects approach, it is also possible to assess whether changes in firearm deaths preceded changes in regulations; evidence of these “lead” effects might suggest reverse causality (i.e., changes in policy in response to changes in firearm deaths). Third, modeling the annual data facilitates flexible modeling of time-varying covariates, including the potential confounding role of other gun policies, which Sharkey and Kang’s estimates do not account for but are likely to be of concern.2,11 Concerns about confounding by, for example, whether a state was controlled by Democrats or Republicans in a particular year could be modeled as a time-varying characteristic rather than reduced to the proportion of years in which the state was led by a Democratic governor over the study period, as it was in the primary analysis, which increases the possibility of measurement error in the potential confounder and residual confounding. WHAT HAVE WE LEARNED? The combination of collapsing all laws into an index and estimating their impact via a change score also complicates our ability to understand the policy relevance of Sharkey and Kang’s analysis. When comparing the results from the change model with the fixed effects specification, which the authors suggest “may allow for stronger causal inferences,” the latter is roughly half the magnitude of the corresponding estimates from the change model. This has important implications for the predicted number of gun deaths averted in each state and the overall interpretation of the results. Furthermore, in reading Sharkey and Kang’s results alongside more recent syntheses of evidence on the impact of gun laws on gun-related outcomes,2 their analysis raises more questions than answers. Some of their estimates are consistent with the literature, such as for the impact of background checks and waiting periods on homicides, but others are divergent. As noted above, the authors estimate little impact of child access laws, yet such laws are considered to have some of the strongest evidence among all gun policies for reducing firearm suicide, particularly among youths. Similarly, Sharkey and Kang report little impact of minimum age of purchasing laws, but the accumulated evidence suggests that these laws are effective in reducing firearm suicides. The authors also estimate little impact of “Concealed carry” laws on any outcomes. This could potentially be explained by including both “shall-issue” and “permitless carry” laws together since the former shows supportive evidence of increasing firearm homicides, but evidence for the latter is inconclusive. In general, it is not possible to reconcile these diverging results since most other analyses have focused on estimating the effects of single policies rather than an index. Sharkey and Kang’s decision to collapse all laws into an index has the merit of simplicity but comes at a steep cost of limiting its value for resolving disagreements about the effectiveness of different gun policies6 and updating evidence syntheses with new results. Finally, to circle back to Sharkey and Kang’s central claim that they show that state gun laws played an important role in reducing gun deaths from 1991 to 2016, it is somewhat at odds with many of the prevailing explanations for why violent crime declined so rapidly across many different geographic areas in the 1990s.12–14 The massive literature on the crime decline makes it clear that there is no single cause, yet it is noteworthy that many of the more promising explanations (changes in age composition, policing, incarceration, lead exposure, community nonprofits, and the decline of the crack epidemic) rarely include changes in gun policy. Sharkey and Kang may hope to alter that perspective with their analysis, but given the challenges outlined here, more analyses of specific laws are likely to be needed to paint a clearer picture. ABOUT THE AUTHORS SAM HARPER and ARIJIT NANDI are Associate Professors of Epidemiology at McGill University. They study the impact of social and economic policies within and between countries on health and health inequalities. They are founding members of the Public Policy and Population Health Observatory (https://www.3po.ca/), a network of researchers studying the effects of policies and programs on health and health inequality using quasi-experimental and experimental methods.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.026
metaresearch head score (Gemma)0.008
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.113
Threshold uncertainty score0.926

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0260.008
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.657
GPT teacher head0.573
Teacher spread0.084 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueEpidemiologySame topicGun Ownership and Violence ResearchFrench-language works237,207