Investigating racial disparities in violence risk assessment using the Spousal Assault Risk Assessment Guide–Version 3 (SARA-V3): Structured professional judgment ratings and recidivism among Indigenous and non-Indigenous individuals.
Bibliographic record
Abstract
Racial disparities in criminal justice outcomes are widely observed. In Canada, such disparities are particularly evident between Indigenous and non-Indigenous persons. The role of formal risk assessment in contributing to such disparities remains a topic of interest to many, but critical analysis has almost exclusively focused on actuarial or statistical risk measures. Recent research suggests that ratings from other common tools, based on the structured professional judgment model, can also demonstrate racial disparities. This study examined risk assessments produced using a widely used structured professional judgment tool, the Spousal Assault Risk Assessment Guide-Version 3, among a sample of 190 individuals with histories of intimate partner violence. We examined the relationships among race, risk factors, summary risk ratings, and recidivism while also investigating whether participants' racial identity influenced the likelihood of incurring formal sanctions for reported violence. Spousal Assault Risk Assessment Guide-Version 3 risk factor totals and summary risk ratings were associated with new violent charges. Indigenous individuals were assessed as demonstrating more risk factors and were more likely to be rated as high risk, even after controlling for summed risk factor totals and prior convictions. They were also more likely to recidivate and to have a history of at least one reported act of violence that did not result in formal sanctions. The results suggest that structured professional judgment guidelines can produce disparate results across racial groups. The disparities observed may reflect genuine differences in the likelihood of recidivism, driven by psychologically meaningful risk factors which have origins in deep-rooted systemic and contextual factors. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".