THE REGRESSION DISCONTINUITY DESIGN: A NOVEL APPROACH TO ASSESSING THE REAL-WORLD EFFECTIVENESS OF HUMAN PAPILLOMAVIRUS (HPV) VACCINATION ON ANOGENITAL WARTS
Bibliographic record
Abstract
Background The human papillomavirus (HPV) vaccine and corresponding vaccination programs have been in Canada for over six years, yet there is no information on their effectiveness in reducing anogenital warts. Objectives To assess the effect of the HPV vaccine (vaccine impact) and Ontario's Grade 8 HPV vaccination program (program impact) on anogenital warts using the Regression Discontinuity Design (RDD; a quasi-experimental, instrumental variable-based technique used to assess the causal effects of policy interventions). Methods By linking Ontario's 36 immunization databases with provincial health records, we will identify a population-based cohort of all girls in Grade 8 in 2003/04-2008/09. Girls will be followed from September 1 of Grade 8 until August 31 of Grade 11. Exposure will be categorized based on HPV vaccine program eligibility (2003/04-2006/07 vs. 2007/08-2008/09) and actual vaccine receipt (0-2 vs. 3 doses). Outcomes will be identified based on a new diagnosis of and treatment for anogenital warts. For the RDD analysis, a continuous instrument will be created using birth month and year, where December 1993 (end of ineligible birth cohort) and January 1994 (start of eligible birth cohort) define the program eligibility cut-off. To estimate the program impact, local linear regression will be employed at the cut-off; here, exposure will be based on program eligibility (intention-to-treat). To estimate vaccine impact, we will use two-stage local linear regression, which will account for actual vaccination status. Several strategies will be used to test and minimize potential confounding bias – e.g., comparison of baseline characteristics, use of triangular kernel, optimal bandwidth selection. Results Based on preliminary data (21 immunization databases), we identified a cohort of 155,999 ineligible and 75,508 eligible girls (N=231,507). Eligibility cohorts were similar across factors like age, urban/rural status, and vaccination history, suggesting they are balanced on factors other than program exposure. A graph of the probability of vaccination by instrument confirmed two additional RDD assumptions – there was discontinuity in exposure at the cut-off (4.6% vs. 45.8%) and no discontinuities at locations other than the cut-off. Conclusions Preliminary results confirm the RDD is appropriate for use in this context. Complete, provincial-level results will be available by June 2013.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.096 | 0.211 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.006 | 0.004 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".