Prediction of violent reoffending in people released from prison in England: External validation study of a risk assessment tool (OxRec)
Bibliographic record
Abstract
We aimed to externally validate the Oxford Risk of Recidivism (OxRec) tool to estimate 1- and 2-year risk of violent reoffending in people released from prison in England. We identified individuals using administrative data shared between official prison and police services. We extracted information on criminal history, clinical and sociodemographic risk predictors, and outcomes. Predictive ability was examined using measures of calibration and discrimination for predetermined risk thresholds. In total, 1770 individuals (median age = 33 [IQR 27-40]; 92% were male) were identified. 31% and 43% reoffended within 1 and 2 years, respectively. Discrimination was good, with AUCs of 0.71 (95% CI: 0.69-0.74) for 1 year and 0.71 (0.68-0.74) for 2-year follow up. At a pre-specified threshold of 40% for 2-year risk, sensitivity was 77% (74%-80%), specificity 54% (51%-58%), PPV 56% (53%-59%) and NPV 76% (73%-79%). Simple model validation found a systematic underestimation of the probability of reoffending. However, after updating the model, calibration was good. External validations of risk assessment tools can be conducted using linked data between prison and police, and may require recalibration before implementation. In this validation, OxRec had good performance on discrimination and calibration measures. It can be considered to be used to improve decision-making about risk of serious offending and the allocation of resources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".