An Examination of the Professional Override in the Level of Service Inventory-Ontario Revision (LSI-OR)
Bibliographic record
Abstract
Despite the overwhelming amount of research conducted on forensic risk assessments in the last twenty years there has been a distinct lack of information on the use of the professional override to adjust actuarial scores. The current study was designed to fill the gap in the research literature examining the effects from using the professional override in the Level of Service Inventory – Ontario Revision (LSI-OR). While there has been recent research conducted indicating that overrides or adjusted actuarial risk assessments are not as accurate as purely actuarial methods (Gore, 2007; Hanson et al., 2007; Hogg, 2011; Wormith, Hogg, & Guzzo, 2012) there is a lack of research conducted solely on the use of professional overrides in forensic risk assessment. This study analysed data from 40,539 provincial offenders in Ontario, Canada. The sample was primarily male (83.9%), White (63.0%), and was comprised of violent (53.0%), sexual (3.3%), and non-violent offenders (43.7%). Predictive validity analyses were conducted to determine the effects of the override for the total sample and then stratified by gender and ethnicity. Special attention was paid to the effects of the override compared between violent, sexual, and non-violent offenders. Results showed that the General Risk/Need score was most strongly correlated with non-violent recidivism over violent and sexual recidivism and that the General Risk/Need was significantly more correlated with non-violent recidivism for female offenders compared to male offenders. Correlation analyses showed that the initial risk levels appeared to be better predictors of general, violent, and non-violent recidivism whereas the final risk levels appeared to be better predictors of sexual recidivism in some cases. For violent and sexual offenders, the initial risk levels were significantly stronger predictors of general, violent, and non-violent recidivism than the final risk levels yet the final risk levels were non-significantly stronger predictors of sexual recidivism. There were no significant differences between the initial and final risk levels’ prediction estimates of the recidivism outcomes for non-violent offenders. Further, there were many more overrides used to increase risk levels than to decrease risk levels overall; sexual offenders had more overrides used to increase risk levels than violent and non-violent offenders combined. Risk level matrices indicated that there were many discrepancies between the number of offenders overridden and their corresponding recidivism rates. Regression analyses indicated additional discrepancies between the significant predictors of recidivism and the significant predictors of the override. Though there were certain methodological limitations to the current study the results still provide important information on the use of the override in a sample of male and female Ontario offenders. The results showed that the override resulted in decreased predictive validity of multiple recidivism outcomes. The conflicting information between the prediction of sexual recidivism and general, violent, or non-violent recidivism prevents a clear message being drawn from this study, yet the equivocal results provide further doubt and criticism of the use of adjusted actuarial practices in forensic risk assessment. Training assessors for how to use the override and examinations of the effects of the override for various offender groups must be improved and more frequently monitored. Further research should also focus on the reasons why overrides are used and if there are any biases concerning certain offender types. Misuse of the override has far-reaching ethical and legal implications that must be limited to ensure the future of forensic risk assessment is as accurate and appropriate as possible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.038 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".