‘Policing Schools’ Strategies: A Review of the Evaluation Evidence
Bibliographic record
Abstract
Background: Schools experience a wide range of crime and disorder, victimizing students and staff, and undermining attempts to create a safe and orderly environment for student learning. Police have long established programs with schools, but there has been no systematic review of evaluations of these programs, outside of police-led prevention classroom curriculum programs such as D.A.R.E. Purpose: This paper documents a systematic search to identify experimental and quasi-experimental evaluations that assess the effectiveness of non-educational policing strategies and programs in schools. Setting: Included studies took place in or around K-12 schools in the United States, Canada, and the United Kingdom. Intervention: Studies were included if they reported on a specific school-based strategy that heavily involved police and did not exclusively involve the police teaching a curriculum or program such as Drug Abuse Resistance Education (D.A.R.E.). Research Design: Systematic review of experimental or quasi-experimental evaluations Data Collection and Analysis: Only those impact studies that used experimental or quasi-experimental design, had at least one outcome measure of school crime or disorder, and were available through December 2009 were eligible. Electronic searches and other methods were used to identify published and unpublished evaluation reports. Findings: The searches identified a total of eleven quasi-experimental studies. Ten of the eleven studies would likely have received a “3” on the Maryland Scientific Methods Rating Scale, a common approach to classifying studies on the basis of internal validity. If evidence rating criteria from the U.S. Department of Education’s What Works Clearinghouse (WWC) were applied, only one study would likely receive a grade of “Level 2” evidence (acceptable with reservations) and the other ten studies would likely not meet WWC evidence screening criteria.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.044 | 0.139 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.008 | 0.006 |
| Bibliometrics | 0.021 | 0.018 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".