Risk assessment and recidivism among Indigenous and non-Indigenous persons: A meta-analysis.
Bibliographic record
Abstract
OBJECTIVE: Risk assessment measures are commonly used in forensic and criminal justice settings to evaluate risk of future recidivism. The use of these measures among Indigenous persons has been the subject of clinical, professional, and legal interest. We sought to add to the literature by examining three frequently studied and clinically used risk assessment instruments among Indigenous and non-Indigenous adolescents and adults in international settings: the Hare Psychopathy Checklist scales, the Level of Service Scales, and the Structured Assessment of Violence Risk in Youth. HYPOTHESES: We hypothesized that the established risk assessment measures would predict reoffending among Indigenous samples and would do so at magnitudes comparable to those of non-Indigenous (White majority) samples. METHOD: We conducted a series of meta-analyses involving three risk assessment measures and indices of general and violent recidivism. We increased the numbers of studies available, particularly among adolescents, by soliciting researchers and reanalyzing data sets. RESULTS: s were .28 and .29, respectively). There was considerable heterogeneity in the magnitudes of effect sizes. Results were not always consistent across age groups, genders, and North American and Australasian samples, and this was particularly true for combinations of variables. There was some evidence that Level of Service Scale Total scores may be associated with lower predictive validity coefficients among Indigenous persons than among non-Indigenous persons. CONCLUSIONS: We discuss the importance of refining and improving risk assessment measures. Findings should be appropriately qualified and interpreted in ways that recognize the impacts of broader sociohistorical contexts on current behaviors. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.014 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".