MétaCan
Menu
Back to cohort
Record W4417093904 · doi:10.1002/cl2.70080

Reliability and Validity of Risk Assessment Tools for Violent Extremism: A Systematic Review

2025· article· en· W4417093904 on OpenAlexafffund
Sébastien Brouillette‐Alarie, Ghayda Hassan, Wynnpaul Varela, Emmanuel Danis, Sarah Ousman, Pablo Madriaza, Inga Lisa Pauls, Deniz Kilinc, David Pickup, Robert Pelzer, Eugene Borokhovski

Bibliographic record

VenueCampbell Systematic Reviews · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicTerrorism, Counterterrorism, and Political Violence
Canadian institutionsUniversité du Québec à Trois-RivièresUniversité du Québec à MontréalConcordia UniversityUniversité de Montréal
FundersPublic Safety Canada
KeywordsReliability (semiconductor)Risk assessmentRisk management toolsValidityGermanField (mathematics)Inclusion (mineral)Predictive validity

Abstract

fetched live from OpenAlex

ABSTRACT Assessment of the risk of engaging in a violent radicalization/extremism trajectory has evolved quickly in the last 10 years. Guided by what has been achieved in psychology and criminology, scholars from the field of preventing violent extremism (PVE) have tried to import key lessons from violence risk assessment and management, while bearing in mind the idiosyncrasies of their particular field. However, risk tools that have been developed in the PVE space are relatively recent, and questions remain as to their level of psychometric validation. Namely, do these tools consistently and accurately assess risk of violent extremist acting out? To answer this question, we systematically reviewed evidence on the reliability and validity of violent extremism risk tools. The main objective of this review was to gather, critically appraise, and synthesize evidence regarding the appropriateness and utility of such tools, as validated with specific populations and contexts. Searches covered studies published up to December 31, 2021. They were performed in English and German across 17 databases, 45 repositories, Google, other literature reviews on violent extremism risk assessment, and references of included studies. Studies in all languages were eligible for inclusion in the review. We included studies with primary data resulting from the quantitative examination of the reliability and validity of tools used to assess the risk of violent extremism. Only tools usable by practitioners and intended to assess an individual's risk were eligible. We did not impose any restrictions on study design, type, method, or population. We followed standard methodological procedures outlined by the Campbell Collaboration for data extraction and analysis. Risk of bias was assessed using a modified version of the COSMIN checklist, and data were synthesized through meta‐analysis when possible. Otherwise, narrative synthesis was used to aggregate the results. Among the 10,859 records found, 19 manuscripts comprising 20 eligible studies were included in the review. These studies focused on the Terrorist Radicalization Assessment Protocol (TRAP‐18), the Extremism Risk Guidance Factors (ERG22+), the Multi‐Level Guidelines (MLG‐V2), the Identifying Vulnerable People guidance (IVP guidance), and the Violent Extremism Risk Assessment (VERA)—all structured professional judgment tools—as well as Der Screener—Islamismus , an actuarial scale. Studies mostly involved adult male participants susceptible to violent extremism ( N = 1106; M = 58.21; SD = 55.14). The types of extremist ideologies endorsed by participants varied, and the same was true for ethnicity and country/continent of provenance. Encouraging results were found concerning the inter‐rater agreement of scales in research contexts (kappas between 0.76 and 0.93), but one of the two studies that examined it in a field setting obtained disappointing results (kappas ranging between of 0.47 and 0.80). Content validity studies indicated that PVE risk tools adequately cover the risk factors and offending processes of individuals who go on to commit extremist violence. Construct validity analyses were few and far between, with results indicating that empirical divisions of scales did not match their conceptual divisions. The internal consistency of subscales was lackluster (Cronbach's alphas between 0.19 and 0.85), whereas full scales demonstrated acceptable internal consistency when assessed (0.80 for the ERG22+ and 0.64 for the IVP guidance). Only one study examined convergent validity, and it revealed a lack of convergence, primarily due to particularities of the scale under study (the MLG‐V2). Discriminant validity analyses were exploratory in nature, but suggested that PVE risk tools might not be ideology‐specific and may apply to both group and lone actors. Finally, although the TRAP‐18 showed a relatively strong postdictive effect size (pooled r = 0.62 [0.35–0.77], p = 0.000), the results were highly heterogeneous ( I 2 = 86%), and all studies used retrospective designs, meaning the outcome was already known at the time of assessment. As such, no included study evaluated true predictive validity (i.e., the ability to forecast future violent extremist outcomes based on prospective risk assessment). This represents a significant evidence gap. Threats to validity were substantial: (a) Many studies were case studies or had very small samples, (b) nearly all samples were constituted through the triangulation of publicly available data, and (c) convenience outcome measures were often used. Although having imperfect data is better than having no data, the current state of empirical validation precludes the recommendation of one tool over another for specific populations and contexts, and calls for higher‐quality validation studies for PVE risk assessment tools. Nevertheless, these tools constitute useful checklists of relevant risk and protective factors that could be taken into account by evaluators who wish to assess the risk of violent extremism and identify intervention targets.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.020
metaresearch head score (Gemma)0.034
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.795
Threshold uncertainty score0.974

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0200.034
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0030.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.078
GPT teacher head0.398
Teacher spread0.320 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSystematic review
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueCampbell Systematic ReviewsSame topicTerrorism, Counterterrorism, and Political ViolenceFrench-language works237,207