How Reliable Are Psychopathy Checklist-Revised Scores in Applied Settings? A Replication and Extension
Bibliographic record
Abstract
The Psychopathy Checklist-Revised (PCL-R) has been described by some as the “gold standard” for assessing psychopathy and is generally accepted as a “reliable and valid” assessment tool. Despite its widespread use, a growing body of research suggests the PCL-R may not be particularly reliable in applied and adversarial settings. Studies suggest the interrater reliability of the PCL-R in field settings (intraclass correlation coefficients [ICCs] ranging .33 to .59) may be lower than posited in the instrument’s manual (ICC = .86 and higher); however, a large portion of these studies were conducted within Sexually Violent Predators evaluations, had small sample sizes, or both. The current study conducted a widescale case law review examining the interrater reliability of PCL-R scores in Canadian criminal justice proceedings. Additionally, potentially differing levels of reliability were evaluated as it pertains to variables such as retaining side, case type (sexual offense versus non-sexual offenses), independence of PCL-R ratings, and demographic information (gender, race/ethnicity, and province). \n\nA total of 176 cases were identified to have multiple PCL-R scores from different examiners. The single-rater ICC was .61 for the total sample, suggesting that nearly 40% of variance was due to error. ICC values were higher for sexual offense cases (.72) than non-sexual cases (.54), suggesting that issues of interrater reliability were not localized to SVP or sexual offending cases. There was evidence of adversarial allegiance such that experts from opposing sides produced lower ICC values (.60-.77), with Crown-retained experts consistently producing higher scores. Independently conducted evaluations demonstrated greater rater agreement (.68) than examiners who were aware of other PCL-R scores (.43), suggesting that awareness of other scores did not increase reliability but in fact lowered it. Rater agreement for Indigenous defendants (.45) was lower than non-Indigenous defendants (.64), suggesting that PCL-R scores may be more reliable for those reflective of the instrument’s early validation samples (i.e., Caucasian samples). Taken together, the PCL-R's low reliability raises concerns about its use, particularly in high-stakes situations. It is recommended that legal/clinical decision-makers consider limitations of the PCL-R and ensure that interpretations are made within the scope of the instrument’s applied reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".