Notice bibliographique
Résumé
Regression discontinuity works in situations where a threshold in a continuous variable is used to assign treatment.1 An example of such a rule would be assigning treatment only to participants with a systolic blood pressure (SBP) of 140 mmHg or higher. Because the probability of treatment changes discontinuously from zero below 140 mmHg to one at and above 140 mmHg, a discontinuity, or abrupt change, should also be observed in the outcome when plotted against blood pressure if treatment has a causal effect. Although the statistical machinery used to analyze regression discontinuity designs is sometimes complex, the basic task is to estimate this discontinuity in the outcome. Many have written that regression discontinuity is the observational design that most resembles a randomized control trial (RCT).2 In both designs, we have precise knowledge of the assignment mechanism. RCTs assign treatment through random allocation and regression discontinuity assigns treatment using a threshold in a measured, continuous variable often referred to as an assignment variable. Therefore, the treated and untreated in an RCT are unconditionally exchangeable, whereas in a regression discontinuity design the treated and untreated are only exchangeable conditional on the assignment variable.2 The added wrinkle of the regression discontinuity assignment mechanism is that there is no overlap in the assignment variable in the treated and untreated groups meaning there is no way to condition on the assignment variable without making important, often impossible, assumptions.3 This is the reason causal estimates in a regression discontinuity design are made at the threshold: because observations close to the threshold on either side are considered to be exchangeable. Although the regression discontinuity design and RCTs share these similarities, we are aware of only one previous attempt to compare the two in a real world setting.4 In this issue of EPIDEMIOLOGY, van Leeuwen and colleagues5 attempt to compare the real world validity and statistical efficiency of the RCT and regression discontinuity designs using both simulations and RCT data. Similar to the previous attempt to compare RCT and regression discontinuity designs, the authors restrict data from an RCT to make it appear as though it were generated from a regression discontinuity design.4 They select a baseline measurement as an assignment variable, create an assignment rule, and drop all observations that do not comply with this rule. For example, the authors use data from an RCT evaluating the effect of the nurse-led intervention targeting cardiovascular risks on the 6-year risk of dementia. They apply an assignment rule where only participants with SBP ≥ 140 mmHg received treatment by eliminating untreated observations with SBP ≥ 140 mmHg and treated observations with SBP < 140 mmHg. This is a clever way of creating a regression discontinuity dataset that should in theory yield causal estimates similar to the RCT estimates. COMPARING RCTS AND REGRESSION DISCONTINUITIES: EASIER SAID THAN DONE Even with comparable RCT and regression discontinuity datasets in hand, however, comparing RCT and regression discontinuity estimates is not straightforward. RCTs estimate an average treatment effect and the regression discontinuity designs generally estimate a local average treatment effect among participants at the threshold. Only under specific circumstances do they estimate the same quantity. The simplest case is when the assignment variable does not modify the treatment effect. In other words, the treatment effect is the same regardless of the baseline value of the assignment variable, in which case the local average treatment effect is equal to the average treatment effect. One test of whether or not this can be assumed is whether the slopes on each side of the cutoff are parallel.3 In none of the examples presented by van Leeuwen and colleagues Figure 1 in reference 5 does this seem to be the case. To the contrary, in all cases, we observe different treatment effects by assignment variable. Therefore, if the authors wish to make a direct comparison, they must find a way to estimate the local average treatment effect with RCT data or to estimate the average treatment effect with the regression discontinuity data. Estimating a local average treatment effect with RCT data can be done relatively straightforwardly by separately modeling the assignment variable–outcome relationship in the treated and untreated groups and taking the difference at the threshold selected for the regression discontinuity estimate. The authors claim to choose this approach when using blood pressure and cholesterol as the assignment variable but only report one estimate. It is not clear to which threshold this estimate corresponds. It is much more difficult to estimate an average treatment effect within the regression discontinuity design when the assignment variable is an effect measure modifier. To do so, one would have to feel confident modeling the relationship between the assignment variable and the outcome even where it is not observed.6,7 In other words, the model for the assignment variable–outcome relationship in both the treated and untreated groups would have to be extrapolated to the side of the cutoff where they were not observed. To demonstrate how difficult an assumption this is, consider the four graphs in Figure 1 in reference 5 and, using your fingers or piece of paper, cover up the lines that would not be observed in the regression discontinuity design (i.e., remove the treated below the threshold and the untreated above the threshold). Of the four graphs, for how many would it be possible to guess or model the unobserved lines? In our opinion, only in Figure 1B in reference 5 would extrapolation correctly give the average treatment effect. But in a regression discontinuity setting, we would not have the benefit of knowing when we were right or wrong. The authors do not use this admittedly difficult method. Instead they claim that by using all the data they are necessarily estimating an average treatment effect. Unfortunately for the authors, the proportion of the data being used does not determine whether one is estimating an average treatment effect or a local average treatment effect. A model using all the data and the appropriate interaction term between treatment and the assignment variable would estimate a local average treatment effect just as a local linear regression model can estimate an average treatment effect where there is no assignment variable–treatment interaction. Properly modeling the assignment variable–outcome relationship is essential in the regression discontinuity design to validly estimate the discontinuity. This is why flexible or nonparametric models are preferred and why interactions between treatment and the assignment variable are most often necessary. Even if one prefers to estimate the average treatment effect, simply leaving out the assignment variable–treatment interaction or using all the data as the authors have done will not accomplish this. Rather, the absence of the interaction can only be justified by homogeneity observed in the data. The end result is that if the treatment effect is heterogeneous across the assignment variable, no realistic modeling strategy can recover the average treatment effect. This is because any regression discontinuity estimate, whether it uses all or part of the data or whether or not an interaction term is included, will be a function of the threshold selected. Therefore, it cannot be an estimate of the average treatment effect, which is independent of the assignment variable. THE POTENTIAL FOR A REGRESSION DISCONTINUITY PROSPECTIVE DESIGN A regression discontinuity design has two important potential advantages over an RCT design, both mentioned in the article: it is potentially easier to recruit participants and it is possible to assign treatment to those at higher risk of the outcome. Both of these advantages may potentially make regression discontinuity designs easier through increased sample size and a lower equipoise standard, respectively. There are, however, also three important drawbacks of a prospective regression discontinuity design. The first is one of the main goals of the article and one that the authors convincingly demonstrate: the regression discontinuity design is less efficient statistically than an RCT design. That this is the case is not in itself surprising, but what is surprising is the magnitude by which regression discontinuity is less efficient. If the important strength of the prospective regression discontinuity design is that it facilitates recruitment, it would have to be anywhere from six to nine times more efficient to achieve the same precision as an RCT design. Considering the additional drawbacks of a prospective regression discontinuity design, one would hope to recruit even more than that. This insight is an important contribution to the regression discontinuity literature. The second drawback is, although a prospective regression discontinuity design could not suffer from nonadherence as would occur in an RCT, it can suffer from a similar phenomenon called assignment variable manipulation (which the authors refer to as selection bias near the threshold value). This problem occurs when people aware of the treatment assignment rule manipulate their assignment variable value to receive their desired treatment. For instance, a doctor faced with an old, overweight man with a SBP of 138 mmHg may choose to record the SBP as 140 mmHg to ensure he receives treatment. If this practice is common enough, many patients with a higher risk of mortality will be shifted from just below the threshold to just above the threshold, likely increasing mortality just above the threshold and decreasing it just below, and biasing the regression discontinuity estimate. The authors say that this type of behavior would have to be avoided in a prospective design, but given the efforts that are made to minimize nonadherence without ever being able to eliminate it in an RCT, one might wonder to whether it is realistic to think assignment variable manipulation could be avoided. There are, however, potential ways of addressing this in a prospective study, most notably blinding doctors and participants to the treatment threshold or assigning treatment after adding random measurement error to the assignment variable. Finally, and most importantly, in any prospective regression discontinuity design, one would never know whether it is possible to estimate an average treatment effect a priori or even post hoc. The assumptions required to estimate the average treatment effect are not verifiable within the regression discontinuity design. This is why it has been suggested that average treatment effect estimates from regression discontinuity designs should only be presented secondary to local average treatment effect estimates and never as the primary parameter of interest.2 Because these authors had access to the entire dataset, these assumptions could be verified. Although the authors estimate the average treatment effect when using age as the assignment variable, one could argue that in all four cases the data do not support estimation of the average treatment effect. It is clear from Figure 1 in van Leeuwen et al. that the effect estimate varies by value of the assignment variable, meaning that a properly modeled regression discontinuity estimate should only be interpreted as a local average treatment effect. One may wonder how frequently it is the case that the treatment effect is homogeneous across the entire range of the assignment variable, particularly when it is an important prognostic variable. This is important because it may be harder to make policy decisions with a local average treatment effect instead of an average treatment effect because one only knows the effect in people near the threshold. Nonetheless, it may sometimes be possible to infer knowledge about this heterogeneity that could be useful in policymaking. For example, if the assignment variable is a strong prognostic variable, treatment effects will be stronger among people with a stronger risk of the outcome. In light of all the advantages and disadvantages of using a regression discontinuity design over an RCT design, we propose that the decision should rest on the lower equipoise standard rather than on any statistical or causal inference argument. It is true that recruitment may be easier in the regression discontinuity design than in an RCT, but to what benefit? The authors demonstrate that the increased number of patients required in a regression discontinuity design would likely make the costs of such a trial prohibitive. As well, although the authors argue that the ease of recruitment would make regression discontinuity studies more representative, this is not necessarily the case. It could be true that better recruitment would make an RCT more representative, but a regression discontinuity analysis, in most cases, would be limited to estimates of the local average treatment effect. Therefore, the contrast should be thought of as a possibly less representative average treatment effect from an RCT against a comparable global effect from a regression discontinuity that cannot be guaranteed and relies on unverifiable and often implausible assumptions. ABOUT THE AUTHORS JEREMY LABRECQUE is a PhD candidate with an interest in quasi-experimental designs and program evaluation. JAY KAUFMAN is a Professor and Canada Research Chair in Health Disparities with interests in epidemiology methods and social epidemiology. Both are based at the Department of Epidemiology, Biostatistics and Occupational Health at McGill University.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,006 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».