How Reliable Are Psychopathy Checklist-Revised Scores in Applied Settings? A Replication and Extension
Notice bibliographique
Résumé
The Psychopathy Checklist-Revised (PCL-R) has been described by some as the “gold standard” for assessing psychopathy and is generally accepted as a “reliable and valid” assessment tool. Despite its widespread use, a growing body of research suggests the PCL-R may not be particularly reliable in applied and adversarial settings. Studies suggest the interrater reliability of the PCL-R in field settings (intraclass correlation coefficients [ICCs] ranging .33 to .59) may be lower than posited in the instrument’s manual (ICC = .86 and higher); however, a large portion of these studies were conducted within Sexually Violent Predators evaluations, had small sample sizes, or both. The current study conducted a widescale case law review examining the interrater reliability of PCL-R scores in Canadian criminal justice proceedings. Additionally, potentially differing levels of reliability were evaluated as it pertains to variables such as retaining side, case type (sexual offense versus non-sexual offenses), independence of PCL-R ratings, and demographic information (gender, race/ethnicity, and province). \n\nA total of 176 cases were identified to have multiple PCL-R scores from different examiners. The single-rater ICC was .61 for the total sample, suggesting that nearly 40% of variance was due to error. ICC values were higher for sexual offense cases (.72) than non-sexual cases (.54), suggesting that issues of interrater reliability were not localized to SVP or sexual offending cases. There was evidence of adversarial allegiance such that experts from opposing sides produced lower ICC values (.60-.77), with Crown-retained experts consistently producing higher scores. Independently conducted evaluations demonstrated greater rater agreement (.68) than examiners who were aware of other PCL-R scores (.43), suggesting that awareness of other scores did not increase reliability but in fact lowered it. Rater agreement for Indigenous defendants (.45) was lower than non-Indigenous defendants (.64), suggesting that PCL-R scores may be more reliable for those reflective of the instrument’s early validation samples (i.e., Caucasian samples). Taken together, the PCL-R's low reliability raises concerns about its use, particularly in high-stakes situations. It is recommended that legal/clinical decision-makers consider limitations of the PCL-R and ensure that interpretations are made within the scope of the instrument’s applied reliability.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».