The Forensic Supplement to the interRAI Mental Health Assessment Instrument: Evaluation and Validation of the Problem Behavior Scale
Bibliographic record
Abstract
Background: Numerous validation studies support the use of the interRAI Mental Health (MH) assessment system for inpatient mental health assessment, triage, treatment planning, and outcome measurement. However, there have been suggestions that the interRAI MH does not include sufficient content relevant to forensic mental health. We address this potential deficiency through the development of a Forensic Supplement (FS) to the interRAI MH system. Using three forensic risk assessment instruments (PCL-R; HCR-20; VRAG) that had a record of independent cross validation in the forensic literature, we identified forensic content domains that were missing in the interRAI MH. We then independently developed items to provide forensic coverage. The resulting FS is a single-page, 19-item supplementary document that can be scored along with the interRAI MH, adding approximately 10–15 min to administration time. We constructed the Problem Behavior Scale (PBS) using 11 items from the interRAI MH and FS. The Developmental Sample, 168 forensic mental health inpatients from two large mental health specialty hospitals, was assessed with both an earlier version of the interRAI MH and FS. This sample also provided us access to scores on the PCL-R, the HCR-20 and the VRAG. To validate our initial findings, we sought additional samples where scoring of the interRAI MH and the FS had been done. The first, the Forensic Sample (N = 587), consisted of forensic inpatients in other mental health units/hospitals. The second, the Correctional Sample (N = 618) was a random, representative sample of inmates in prisons, and the third, the Youth Sample (N = 90) comprised a group of youth in police custody. Results: The PBS ranged from 0 to 11, was positively skewed with most scores below 3, and had good internal consistency (Cronbach's Alpha = 0.80). In a test of concurrent validity, correlations between PBS scores and forensic risk scores were moderate to high (i.e., r with PCL-R Factor two of 0.317; with HCR-20 Clinical of 0.46; and with HCR-20 Risk of 0.39). In a test of convergent validity, we used Binary Logistic Regression to demonstrate that the PBS was related to three negative patient experiences (recent verbal abuse, use of a seclusion room, and failure to attain an unaccompanied leave). For each of these three samples, we conducted the same convergent validity statistical analyses as we had for the Developmental Sample and the earlier findings were replicated. Finally, we examined the relationship between PBS scores and care planning triggers, part of the interRAI systems Clinical Assessment Protocols (CAPs). In all three validity samples, the PBS was significantly related to the following CAPs being triggered: Harm to Others, Interpersonal Conflict, Traumatic Life Events, and Control Interventions. These additional validations generalize our findings across age groups (adult, youth) and across health care and correctional settings. Conclusions: The FS improves the interRAI MH's ability to identify risk for negative patient experiences and assess clinical needs in hospitalized/incarcerated forensic patients. These results generalize across age groups and across health care and correctional settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.063 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".