PROTOCOL: Impacts of after‐school programs on student outcomes
Bibliographic record
Abstract
Nationwide, an estimated 8 million children between the ages of 5 and 14 are frequently unsupervised after school (NIOST, 2003). Recent data reveals that more than two-thirds of low- and moderate-income youth do not have parental supervision available after-school due to parental work requirements (Long & Clark, 1998; U.S. Bureau of Labor Statistics, 2000). These statistics are not surprising given the need for low-income families to meet pubic assistance work requirements, and the correlation between low to moderate income with single parent households. Research has linked such unsupervised time with increased risk-taking behaviors, victimization, and poorer academic outcomes (Dwyer et al., 1990; Newman et al., 2000; Osofsky, 1999; Posner & Vandell, 1999; Richardson et al., 1989; U.S. DHHS, 1995; U.S. DOE & U.S. DOJ, 2000). Unstructured, unsupervised after-school time has increasingly been seen by policy makers and the public as holding “risk and opportunity” (Hofferth, 1995). And, after-school programs have been touted as a means to reduce negative behaviors and improve positive outcomes, especially for lower-income, urban students. Within the last few years, after-school programming has seen tremendous growth. The federal government, states, localities and private foundations have invested substantial money and resources in programs. For example, appropriations for 21st Century Community Learning Centers have increased from $40 million in 1998 to the near $1 billion that is currently appropriated for the program. In this short period of time, the number and strength of advocacy groups in this field has also experienced a great deal of growth. As evidence of their voices, tremendous fervor surrounded the recent release of the first year findings from the national evaluation of 21st Community Learning Centers (CCLCs) (U.S. DOE, 2003). Several criticisms were directed at this report2, but arguably the strong responses to the report's primarily null findings were likely more reactions to the use of a single, high profile experimental study to recommend a 40% reduction in 21st CCLC appropriations. The resulting lobbying and grass roots efforts to maintain or increase the 21st CCLC appropriations served to highlight that continued support for a high investment in and expansion of after-school programming is not supported by a large or strong research base. Several quasi-experimental and non-experimental studies are frequently cited as evidence that after-school programming promotes positive developmental and emotional outcomes in low-income youth, may help to improve academic outcomes, and may decrease student participation in criminal or violent activities (Baker & Witt, 1996; Foley et al., 2000; Huang, et al., 2000; Jones & Offord, 1989; Le and Hamilton, 2001; McLaughlin & Irby, 1994; Posner & Vandell, 1994; Ross et al., 1992; Schinke et al., 2000; U.S. DOE & U.S. DOJ, 2000; Grossman et al., 2002; Welsh, et al., 2002). Some of these studies have compared participants’ outcomes to those of non-participating youth. However, their designs cannot completely control for differences (like motivation) between the youth who volunteered for the program and those who did not. Does this matter? Recent research shows that it does. Strong bias affects the estimates from quasi-experimental studies, especially for studies of voluntary participation in programs (Guyat, et al., 2000; Agodini & Dynarski, 2001; Weisburd, Lum & Petrosino, 2001; Wilson & Lipsey, 2001; Glazerman, Levy & Myers, 2003). For this reason, a comparison of outcomes between program and non-program youth does not accurately measure what the program has added to participants’ development. Further work is being conducted to assess whether any quasi-experimental methods accurately replicate experimental results, and under what conditions (Glazerman, Levy & Myers, 2003). Until the field is more confident in how to control for the possible biases and/or which kinds of quasi-experimental studies, if any, can be reasonably substituted for experimental designs, it seems imprudent to combine reliable with unreliable impact estimates when answering questions on program effectiveness. The field of after-school programming could benefit from a review of high-quality experimental, evaluations of the programs that have experienced the greatest amount of growth since funding for programming increased in order to better understand the current state of the evidence, and the impact these programs may be having on participants’ academic, behavioral, and social/emotional outcomes. The current increase in funding has allowed more districts and schools to offer programs that combine recreational, academic, and youth development programming. These programs are thus likely attended by a majority of youth in organized after-school programming and are therefore currently of great interest to policy makers.3 We identified seven recent, major reviews of research on the impact of out-of-school programs on student outcomes. Three of these reviews focus on programs specifically designed to promote positive youth development (Catalano, et al., 2002; National Research Council, 2001; Roth et al, 1998). The other four reviews have considered programs with additional goals, like improved academic achievement. These four reviews are most relevant to the current interest in after-school programs, and so are discussed below. Review of Extended-Day and After-School Programs and their Effectiveness identified and described programs with an educational focus that had shown some evidence or promise of effectiveness and/or had the potential for dissemination and replicability (Fashola, 1998). Both experimental and quasi-experimental studies (“well-matched treatment and comparison groups”, p.7) that measured achievement and other outcomes were included. Fashola concluded that a number of the programs looked promising, but few of the collected studies had rigorous designs and almost all of the studies suffered from selection bias, limiting the conclusions that could be drawn with confidence. The information and issues raised in the Fashola review were useful as a guide to the state of knowledge in the field at the time the report was written, and to suggest areas for further study. However, its current usefulness for guiding after-school programming is limited for three reasons. First, not all the studies included in the review were implemented in the after school context. Some were programs that existed within the normal school day, although the author noted that they had the potential for replication during the after school hours. Second, the evaluations included mentoring and tutoring programs, which are fundamentally very different than what we will define as more traditional after-school programs. Finally, it is not clear the extent to which the review included studies which showed null or negative impacts. Eccles and Templeton (2002) conducted a review similar to the Fashola (1998) review by casting a broad net in defining after-school programs, including mentoring programs, sexual education programs, or programs with a very low intensity and/or short duration. This review also included both quasi-experimental and experimental design studies. Similar to Fashola, Eccles and Templeton found that the research field in this area is still quite young and inconsistent, citing few experimental studies, little overlap between evaluated outcomes between studies, and the lack of implementation and process data to help understand impact findings. But, they do draw some preliminary conclusions from the experimental and quasi-experimental studies. They suggest that “there is growing evidence that youth programs focused on both prevention and promotion do increase positive outcomes and decrease negative outcomes” (p.172) and that “programs not explicitly focused on academic instruction produce gains in academic achievement, school engagement, and high school graduation rates … (as well as) declines in school-related problem behaviors.” (p.172) Evaluations of After-School Programs: A Meta-Evaluation of Methodologies and Narrative Synthesis of Findings uses similar inclusion criteria to the above referenced reviews (Scott-Little et al., 2002). The review suggests that after-school programs may positively affect standardized test scores and homework completion, and that programs may have more of an effect for younger and more academically “at-risk” students. Because their review included both experimental and quasi-experimental designs, the reviewers do not make statements of causality in these areas. However, using results from two experimental studies, they suggest that programs can have positive impacts on participants’ social/emotional outcomes. However, these results are from programs that are beyond the realm of the traditional program experienced by the majority of youth; we are left questioning the generalizability of such findings. Hollister (2003) reviewed the effectiveness of after-school programs in The Growth in After-School Programs: Status, Issues, and Evaluation Impacts. Unlike prior reviewers, he limited his search to experimental design studies. However, Hollister also cast a broad net in the definition of an “after-school program” (including mentoring, tutoring, remedial schooling, and comprehensive services programs). Looking across the ten experimental design studies identified for review, he found that mentoring and tutoring programs have positively affected in-school and out-of-school outcomes, parent involvement and training has been an effective component, and life-skills training curricula may positively effect some out-of-school outcomes. Because of the diverse characteristics of the included programs, these findings are again difficult to apply to the majority of students currently participating in the more traditional programs which may combine academics with recreation and youth programming. As the authors cited above have noted, their reviews were limited by the amount of available evaluation work in this very quickly growing field. However, other factors related to their inclusion criteria limited the conclusions they were able to draw. Because prior reviews included studies of both experimental and quasi-experimental methodologies, and/or programs which look quite dissimilar in design and service delivery, it has been difficult to use such work to answer questions about program effectiveness. This proposed review will differ from prior reviews in two main ways. First, this review will make use of recently released experimental studies in this field that were not captured by prior reviews. Second, this review will limit the inclusion criteria up front, so that the answers to policy relevant questions, like program effectiveness, can be made with more certainty and drawing upon a more homogeneous set of program models. A systematic approach to the identification, compilation and analysis of well-designed and executed studies measuring the impacts of program participation on student outcomes will allow us to answer these questions as best one can given the available research base, and will also point out the limits of our knowledge. Studies must meet the following criteria to be included in this narrative review and meta-analysis. Programs must operate after-school, and may or may not have a before-school component. They must not be targeted specifically at youth with particular special needs, such as learning disabilities, physical disabilities, emotional problems, or behavioral problems. Summer school programs or programs that have a significant in-school component will not be included. Programs must provide academic, recreational, and/or positive youth development activities. However, the primary means of attempting to achieve positive outcomes in these areas must not be through a one-on-one mentoring or tutoring format. While most of these programs do operate during the after school hours, the design and delivery of such programs assume a much different relationship between program teachers/volunteers and youth, and therefore such programs would fall out of the bounds of this particular review. In other words, programs chosen for this review will fall under the general heading of more traditional after-school programs than some of the programs captured by prior reviews. These more traditional programs have arguably experienced the greatest amount of growth since funding for programming increased, allowing more districts and schools to offer programs. These types of programs are thus likely attended by a majority of youth in organized after-school programming. Such programs are therefore currently of great interest to policymakers. However, we want to make clear that we expect that there will still be some degree of heterogeneity among program models. For example, some programs may offer more recreational than academic activities, while others might focus on academics for most of the program hours and then only offer recreational activities one day a week. Programs could operate in a variety of settings – schools, community centers, and religious institutions. But, the evaluation should report results separately for programs operating in different settings. Finally, we will limit our focus to studies of interventions conducted in North America.4 We will not specifically search for any international studies because the intervention context would likely be much different than that of the U.S. and Canadian programs. However, we will provide a full citation of any international studies we find to aid future reviewers in this field. Participants must include youth enrolled in regular public or private K-12 schools; typically, we expect these youth to range in age from 5–19. This large range was chosen so that studies which span a wide range would not be excluded. Also, some potentially relevant programs are designed to prevent risk-taking behavior for students age 13 and above. In order to be included for consideration in the review, evaluations must use well-implemented experimental designs, and the control group must not have received a similar intervention. Parental consent for study must have been received prior to random assignment. Otherwise, we would assume there to be selection bias among those consenting for study after the random assignment process. Furthermore, description of the methodology must be clear and complete enough that the reviewers can judge the rigor of the implementation of design. Studies should also provide a thorough description of program goals and activities. We will include only studies published in 1982 or later for two reasons. First, Lauver's (2002) prior search did not identify any studies prior to this date. Second, there was little public support for programs and evaluations in this field before the 1990s. We propose to retain for inclusion those studies which meet the criteria outlined above. Then, specific codes designed to judge quality of the design and implementation of the study will be applied. Only those studies that are judged to have a high likelihood of generating unbiased estimates of program impacts will be included in the statistical reporting on estimated program impacts. Dimensions that will be considered when judging study quality are found in Appendix A. We will include studies with academic (test scores, grades), social/emotional (self-confidence, self-esteem, aspiration, locus of control), and/or behavioral outcomes (risk-taking, time use, school attendance). We understand that evaluators may have chosen to measure outcomes that may not have been the specific goals of the programs. However, we don't believe that these outcomes should be excluded from the analyses. Prior research and reviews have used non-experimental research to generate hypotheses about the cross-over between seemingly incompatible program goals and measured outcomes. For example, the review recently conducted by Eccles and Templeton (2002) suggests that programs not explicitly focused on academic instruction can produce gains in achievement, school engagement, and high school graduation rates. Several other recent studies also support the hypothesis that youth development goals can promote academic achievement (National Research Council and Institute of Medicine, 2002). Additionally, many logic models for after-school programs include both academic and social/emotional outcomes, recognizing that the two are likely closely related (HFRP, 2003; Lauver, 2002; Dynarski, et al., 2001). Upon the reporting of results, the reviewers will take care not to suggest that programs which may not have had academic goals should be held accountable for achieving them. Similarly, programs which may not have had behavioral goals should not be held accountable for achieving them. However, it would be of great policy importance to understand what outcomes are being achieved, and from which types of programs. Recent reviews and studies have highlighted the probable certainty of strong bias affecting the estimates from quasi-experimental studies, especially for studies of voluntary participation in programs for children and adolescents like after school programs (Guyat, et al., 2000; Agodini & Dynarski, 2001; Weisburd, Lum & Petrosino, 2001; Wilson & Lipsey, 2001; Glazerman, Levy, & Myers, 2003). However, the field of after-school programming is awash in quasi-experimental studies as funding or feasibility has traditionally limited the possibility of experimental evaluations of this particular intervention. Using these criteria to find references to quasi-experimental studies will facilitate possible future work to explore the sensitivity of impact estimates across studies employing different The field of after-school evaluation is young and quickly after the of 21st Century Community Learning The Research (HFRP, and (2002) have knowledge of the current programs that have or are currently rigorous The Research has recently published an Evaluation (HFRP, 2003). The a program and evaluation description of out-of-school studies of and after-school programs. (2002) for and impact evaluations of after-school programs for low-income youth by an search of major in also the and program evaluators and other in the field to identify studies. Both have to great to identify the available studies in this field and have criteria before including studies in their we will our of studies to for inclusion from these reviews. We will also in the field to us in our search with the most current of the studies that are included in the (2002) review and were not published in but were of the or For this reason, bias may not be a will identify the of the study and the and these particular issues of potential bias will be considered at the time of studies that have been identified by and the prior reviews cited above first as and were not also published in review and a limited of time and we will not search for this review. However, we will the that the search of the major so that future reviews can all after that that are then by the As a of this not by the will not be but knowledge of the field suggests that we will likely not rigorous experimental design studies of after-school programs when drawing such search We will identify studies for inclusion from any prior reviews conducted in this and also by the of other studies identified through the search process. A search using a search such as will be included. We that major research studies could be identified through this prior to the results in Both primary reviewers will review the of all studies, and will which studies should be Only those studies which to meet the above criteria for inclusion will be between the will be by a to any study that is in Both reviewers will the full studies and recommend which studies should be included in the review. between the will be by the to the reviewers, Dimensions that will be considered when are found in Appendix Both reviewers will the first studies. in will then be in order to the of time and resources of any studies, one will then all the other studies, while a will a random of these studies. like will be for all studies. The reviewers will to the authors of studies that are data that is for the review. Because we have drawn very the types of after-school programs and evaluation design which would be included in this review, we that any conducted will include a number of studies. the design and delivery of these programs may look quite and we would not be to it is likely the findings from prior that the evaluations measured different outcomes in different ways. therefore may not be to the of and then the of all available studies. In an to whether after-school programming has a effect in any one we would of like scores, scores, and a study a in we will that study only one data point to the analysis for in order to of the Some after-school programs have had evaluations In these we will the outcomes by the of the period outcomes at the by one two one two years, a number of studies are we will also outcomes of such studies by the in effect across We expect to use by to any However, we also to issues of different reporting and data to support the meta-analysis. We will and for on when and how to and and Wilson for assistance using different reporting We will not be able to identify what effect will be used for the we have identified the outcomes measured and how they have been We to report the effect in a and also report the are likely more and relevant to of reviews. 8 our knowledge of the experimental evaluations of after-school programs prior to the review, we are confident that we will identify a limited number of studies. these studies likely would not the of possible studies that can be on after-school programs that the characteristics of the intervention we for this review. For this reason, we will likely a for all & In order to meet the of this review, data will be an of the The systematic review will include research findings from well-designed and studies, which the impact study findings. and data will provide of program goals, activities, made to the school day, and program context. The data will be used to how and context can the outcomes of after-school programs. The authors will the review upon the release of the year findings from the national evaluation of 21st Century Community Learning Centers being conducted by Then, the review will be as This was with the help of many including the of the of of on systematic reviews and organized by has an impact evaluation of an after-school program random that will be reviewed for its potential inclusion in the systematic review of the will from the to include this research study in the and will this research study and make its
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".