PROTOCOL: Interventions Intended to Reduce Pregnancy‐Related Outcomes Among Adolescents
Bibliographic record
Abstract
The main goal of this review will be to examine the effectiveness of intervention programs in affecting the sexual risk-taking behaviors of adolescents and to explore whether there is evidence to suggest that particular intervention strategies have been effective in the settings in which they have been tried. Studies have shown that teens who become pregnant, and especially those who give birth at a young age, face hardships that can have detrimental economic and social consequences, both in the short-term and as the young mothers make the transition to adulthood (Maynard, 1996; McLanahan, 1994; Moore et al, 1993). While the pregnancy and birth rates for teenagers in the U.S. have declined over the past decade, they remain the highest of all industrialized countries (Alan Guttmacher Institute, 1999). Other industrialized countries such as the United Kingdom, Canada, and Australia also have relatively high rates of teen pregnancies and births (Maynard, 1996). It is therefore important to understand the factors that lead teens to engage in risky sexual behaviors and to examine the effectiveness of interventions that attempt to reduce sexual risk-taking. There is a great deal of research exploring the antecedents of sexual risk-taking behaviors. Researchers have identified a variety of correlates of sexual-risk taking among teens, including community characteristics; school characteristics; family characteristics; biological factors; psychological factors; relationships with peers, parents, and school; as well as attitudes and beliefs concerning sex (Bearman et al., 1999; Blum et al., 2000; and Kirby, 2001). It is clear that myriad factors influence adolescents' decisions to engage in risk-taking behaviors. However, it is not clear what strategies will be effective in curbing these behaviors. Many types of interventions designed to reduce teen sexual activity, prevent teenage pregnancy, and childbearing have been implemented over the course of the past few decades,1 focusing on various combinations of the antecedents identified above. In addition, a multitude of actors have entered the pregnancy prevention arena ranging from schools to community-based organizations to religious organizations (U.S. Department of Health and Human Services, 2000). Policymakers, researchers, and practitioners have engaged in ongoing debates concerning the content and timing of pregnancy prevention programs. Specifically, two major issues have been hotly debated: (1) whether school-based sex education programs should have an abstinence-only focus or whether such programs should also include information and education on contraception protection, and (2) whether pregnancy prevention programs should be aimed at younger versus older adolescents. There have been a number of reviews of teenage pregnancy prevention programs (see Table 1). While there is substantial overlap in the studies included in these prior reviews, these reviews differ in their inclusion criteria, analytic strategies, and their conclusions regarding the effectiveness of interventions. There is considerable inconsistency in terms of the methodological standards reviewers have applied in their inclusion criteria. For example, some include only experimental design research; some include experiments and quasi experiments; and some include even studies with no matched control group (see Table 2). None of these reviews has specified clear criteria as to what constitutes a well-executed study under particular methodology; and few have followed clear and consistent search criteria. While some have conducted statistical meta-analyses, many provide only a narrative description of study findings. Finally, some of these reviews are outdated and others are too narrowly focused. There are two relatively recent prior reviews that attempt to present a comprehensive review of the effect of teen pregnancy prevention interventions warrant brief discussion. One is a narrative and vote-counting review of pregnancy prevention programs in the U.S. and Canada (Kirby 2001). The other is a review of pregnancy prevention programs in North America that follows procedures similar in many ways to the Campbell Collaboration review methods, including conducting a statistical meta-analysis (DiCenso et al. 2002). Interestingly, the two reviews include overlapping but different studies, and they draw different conclusions regarding the effectiveness of intervention strategies. Commissioned by the National Campaign to Prevent Teen Pregnancy, Kirby (2001) includes over 70 primary studies in his update of a 1997 review. This review covers a variety of intervention types, such as sex education programs, school-based clinics, community-wide pregnancy prevention efforts, and youth development programs.2 Kirby examines the effect of programs on a range of outcomes related to sexual behaviors, contraceptive use, pregnancy, and birth rates. His stated study inclusion criteria encompass evaluations that rely on either experimental or quasi-experimental designs. He presents the specific characteristics and findings for each included study in tabular form and then synthesizes the study findings in his narrative presentation by means of ‘vote counting’ – reporting and drawing conclusions based on the number of studies that reported statistically significant positive (favorable) effects, negative effects, or no effects. Kirby (2001) concludes from his review that effective programs are characterized by the inclusion of certain key components. However, these conclusions are based on looking at the common characteristics of the programs that showed evidence of statistically significant impacts. The significant impacts were not necessarily observed across all outcomes. Moreover, it appears that many of the programs that had characteristics of the “effective” models did not show evidence of effectiveness. Unlike the Kirby (2001) review, DiCenso et al. (2002) limited their review to randomized control trials (RCTs).3 This review, which includes both published and unpublished studies, has clear, thorough, and documented inclusion criteria. DiCenso and her colleagues conducted statistical meta-analysis looking both at the overall study findings and at findings for sub-groups defined by the characteristics of the intervention. The review concluded that there were no program types that consistently reduced sexual risk taking sexual behaviors of adolescents. One of the major sources of the difference in the conclusion from this review and that in Kirby (2001) is the fact that Kirby included many more studies, most of which used non-experimental methods. The majority of the significant program impact estimates in the Kirby review were from quasi-experimental studies—a finding that is consistent with a methodological study by Guyatt et al. (2000) that preceded the DiCenso et al. review. The DiCenso et al. review has three main limitations. First, although the review included only RCTs, it included studies regardless of the quality of the study implementation and, indeed, the authors documented nontrivial quality concerns with most of the studies. The authors created a “quality index” that ranges from 0-4 (4 is highest) based on the following four criteria: (1) appropriate randomization, (2) unbiased data collection, (3) a minimum of 80 percent of the sample included in the follow-up outcome data, and (4) less than a 2 percent difference in attrition rates between experimentals and controls. Only 8 of the 26 studies earned a quality score over 2, suggesting that the internal validity of many of the findings may have been compromised. In their meta-analytic results, the authors present pooled effect sizes for all of the included studies and do not differentiate between studies that are more likely to be internally valid versus those that are not. Second, the authors chose only to include studies conducted in the United States and Canada. Finally, the authors only included studies for which it was possible to create separate effect sizes by gender, thereby, excluding potentially internally valid studies that did not provide gender-specific information. The proposed study will improve upon the prior reviews, particularly the Kirby (2001) and DiCenso et al. (2002) reviews, in six ways. First, we will focus the review on a clear and policy relevant set of questions in terms of both the intervention and the outcomes. We will focus on interventions with a primary goal of reducing sexual risk-taking behavior, and studies of interventions that include measures of at least one of three key outcomes measures—(1) sexual initiation, (2) sexual activity and contraceptive use, which we use to construct a measure of pregnancy risk, and/or (3) pregnancy. Second, we include evaluations operating in a broader set of geographical contexts than have most prior reviews, while bounding the search to encompass research on programs that have operated in developed countries with relatively high rates of teen pregnancy. The clarity of the boundaries will make it efficient to augment the review to include a broader set of geographical contexts at a later date. Third, we will limit the prospective research base to those studies with a strong potential for generating credible (internally valid) findings. Specifically, we will include in the search inventory only randomized control trials that address the study questions within the geographic and language boundaries for the search. Fourth, we will evaluate whether the research base is an adequate representation of the programs currently in operation and we will assess the appropriateness of combining effect sizes of different program types. While we do not have the resources to do a systematic and thorough assessment of all programs in operation, we will rely on summary reports by government and non-government entities to inventory strategies aimed at preventing teen pregnancy. We will determine the extent to which there is credible evidence of the impacts of these particular strategies by comparing this inventory with the studies included in our review. Fifth, we will explore differences in outcomes among clusters of programs defined by seemingly important programmatic features such as whether the program has an abstinence-only focus or includes contraception information. Finally, we will note those studies that have been excluded due to study quality considerations and the primary reasons for their exclusion.4 If there a sizeable body of evidence: The following sections detail our proposed approach to the review. Section A details the criteria for inclusion and exclusion of studies in the review. Section B describes the search strategy and defines the boundaries within which the search will be conducted. Section C documents the methods generally used in component studies. Section D provides the criteria for determination of independence of findings. Section E details our strategy for coding information from the studies included in the review. Section F outlines our plans for conducting the statistical analysis of overall program impacts and for exploring estimated impacts for subgroups defined by program qualities and the characteristics of the target population. Finally, Section G discusses the treatment of qualitative research. We propose to include all studies that meet the inclusion criteria outlined above, regardless of publication status, and we will code basic information on all studies regardless of the study quality. Our search strategy will make use of electronic data bases, hand searching of journals, internet searches, and personal contacts. For each search, we will maintain a log documenting our procedures and their yield. While our intention is to unearth all relevant studies, time and financial constraints prevent us from conducting extensive hand searches and database extraction. However, we believe that our search plan will uncover nearly all relevant studies, and the transparency of our plan will allow future researchers easily to expand their searches beyond our specified boundaries. Based on preliminary keyword searches, we believe that these general labels will capture the majority of the studies. The remainder of the studies will be accessible through alternative measures of literature searching, as mentioned below. Abstracts will be collected for all seemingly relevant studies. If the abstract appears appropriate, then the full study will be obtained and reviewed. For more information concerning study inclusion, see Section 2: “Methods of review” and Section 3: “Selection of trials,” below. Due to timing, financial, and technology constraints, all databases will be searched for documents written or published between January 1, 1992, and December 31, 2002. Many of the databases are not fully updated and electronic searching is not fool-proof. Thus, we propose to conduct a hand search of the past 10 years in the following journals that are highly likely to contain relevant studies: AIDS Education and Prevention, Journal of Adolescent Health; Pregnancy Prevention and Youth; Journal of Adolescent Research; American Journal of Public Health; Journal of Health and Social Behavior; Journal of Sex Research; Family Planning Perspectives. We will add journals to hand-search if we find, through our database searches, that other journals have highly relevant studies. In addition to conducting key-word searches in the on-line databases, we also will hand search those journals noted above for the years 1992 through present. In light of the number of prior reviews in this area, our financial constraints, and our own familiarity with the research, we are quite confident that hand searching journals prior to the last 10 years will yield few to no new studies. By having a clear “stop date” for our hand searching, others who feel we may have missed important studies through this restriction will be able to conduct a complementary search of earlier journal issues. All relevant government, foundation, professional associations and policy research firm websites will be searched. In addition, keyword searches (see the above list of keywords) will be conducted using search engines such as google.com. Personal contacts: An initial library of studies will be assembled from Doug Kirby's collection. In addition, principal investigators in a current ongoing evaluation of abstinence-based programs will be consulted. Reference lists: We will check reference lists of review papers and of primary studies to uncover additional studies that may be eligible for inclusion. Search log: A comprehensive log will be maintained that will keep track of all databases searches, key words used, number of hits, and time spent searching. Full citations on all relevant or potentially relevant studies will be maintained in the review database. As we identify studies that potentially are relevant to the review, we will systematically gather details on the study (progressing from title review, to abstract review, to full study review) until we are able to determine with certainty whether the study is or is not appropriate for inclusion. The initial step is to review titles and abstracts to rule out obviously inappropriate studies. The second step is to collect the full reports for all potentially eligible studies and to assess the appropriateness of the study for inclusion in the review based on reading the full report. Two reviewers (identified by their reviewer code) will do all screening independently. Differences will be resolved by negotiation, including a third party if necessary. In some instances, the review team will find it necessary to contact the study authors to gather additional information. Such contacts and the supplemental information gathered for the review will be documented. The review will include bibliographical information on all potentially relevant studies identified through the screening process. However, only the subset of studies meeting a base set of criteria required to generate credible evidence on program impacts will be included in the descriptive and statistical analyses and reporting of study findings. As we review each study, we will code the basic information needed to determine whether the study meets the overall inclusion criteria. For those studies that meet inclusion criteria, we also will code the information required to determine quality of study implementation and the information required for the analysis of program impacts for the total sample and for key subgroups/ program types. Ultimately, we will include in the formal analysis estimating program impacts only those RCT studies that meet the criteria noted above. We will analyze results for the various outcome measures— sexual initiation, pregnancy risk, and pregnancy— separately. We will never pool effect size estimates from effect sizes within a study will be included in the following (1) if a study reports effect sizes by (2) if separate outcomes are for different within a and/or (3) if separate effect sizes are across different study within a In there are of follow-up for a outcome we will focus on the follow-up for which there is adequate of the study sample least percent of the initial study If there are an adequate number of studies with we will examine in effect sizes over we will and the of systematic in effect sizes systematically over In we studies with outcomes for or overlapping one control we will code all of the effect sizes but only include one in the We will which treatment to include in both are within the to be programs and their will be as study will be and potentially used in the analysis as a descriptive All studies meeting the initial criteria will be using an that four main sections (see Section A includes all relevant bibliographical information. Section B basic information concerning both the study design and intervention criteria. Specifically, this includes questions the research design whether there was adequate follow-up of the whether the program the group of the geographic of the the primary goal of the and the types of outcomes If all of the criteria in Section B are then C and D are Section C information concerning the and implementation of the and Section D additional information concerning study sample and outcome data needed to effect We will use to code and our database. such a database us to or in to our coding Moreover, coding created using are and, easily among Two will code the studies. will differences in coding decisions and coding this the primary will code all of the studies, and a second will code a sample of percent of the studies. If there are more than 10 percent in between the two in the the 80 percent of the studies will be by a second and all differences in coding This is based on our prior high of among who have and are in systematic review methods. analyses will be conducted in If statistical meta-analysis is appropriate, analyses will be conducted using and to the it is necessary that all outcomes are as effect or of and sample The data we are able to from the evaluation reports are to in their form and in their to be of these we will make necessary and to the data using by 2001). All three outcome measures are and, will be reported as and For studies information is to an effect but study criteria are we will code the of the effect for descriptive and to and/or findings from the We will not use to In study impact these will form the for our impact We will not for of the outcomes. However, we will some we have information on sample but have information to sizes For example, a study may provide sample sizes for program and control sample attrition and follow-up but no information on follow-up program and control sample In such a we use the information to create estimates of the total number of program and control group in the follow-up sample and the number who have the outcome If there are we have adequate of studies different program qualities and implementation we will conduct a analysis to explore whether it is to the studies of various sample and program The analysis will be conducted The findings of such a analysis will be descriptive of the differences in impact across the subgroups in the studies. Differences between subgroups be as evidence of relationships between the and the of program impacts. they may for regarding program effectiveness that be through studies. types of measures will be and potentially in the statistical (1) the treatment control group no treatment versus (2) the intervention sex education programs, based or (3) qualities of the intervention of whether the intervention was (4) the target for the intervention high risk versus risk school versus high school of randomization, of these has been identified in models as potentially important in program impacts. We likely will not have more than studies in our meta-analysis and, will be limited in terms of the number of subgroups we can examine at one However, by the information on all of these it we will the for future such analyses if the number of studies The analysis will be conducted using appropriate, we will use to for sample analysis will be conducted to determine whether effect sizes due to the following (1) for a similar outcome example, is in a variety of ways across (2) comparing results by study sample attrition (3) different estimates there is not information to for in randomized If additional analyses will be conducted if other methodological issues that may our in the estimated pooled effect size It is that there will be considerable among the types of and settings for the programs by the evidence models likely will be most will be conducted to whether to use a or from such will be reported in the review. If estimates are statistically significant and it is not possible to in the pooled effect size using a few basic study then pooled estimates will be based on We will code key qualitative information the program intervention and study This information will be used for descriptive This review, will be updated two years to include additional study The primary will the lead in this review. This was with the of many colleagues Doug Kirby, and of the of Education on conducting systematic reviews, which included the following and has of related to this review. experimental design studies in this However, of her studies questions that for for this review.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".