Immigration Detention and Release Decisions in Canada: Development and Preliminary Validation of a Risk Assessment Tool for Frontline Officers
Bibliographic record
Abstract
Immigration detention systems face mounting pressure to demonstrate transparent and defensible decision-making practices amid growing ethical concerns. The Canada Border Services Agency (CBSA) has drawn particular scrutiny in this area, largely due to its partnerships with correctional agencies. Critics challenge CBSA's framework for placing noncitizens in these facilities, characterizing its risk assessment processes as opaque and arbitrary. To address these limitations, we developed the Immigration Risk Assessment for Detention (IRAD), an empirically informed tool designed to meet CBSA's multiple decision-making needs—from release on community-based alternatives to detention (ATDs) to security classification level within detention facilities. This dissertation presents research conducted across three co-authored articles, each representing a distinct phase in the IRAD’s development and preliminary validation. First, we surveyed 92 CBSA employees to gather their insights on immigration detention risk assessment. Second, we developed a 30-item numerical IRAD prototype by integrating our survey results with CBSA’s operational guidance and correctional risk assessment research. We then conducted a longitudinal retrospective validation of the IRAD prototype using 301 case files, which provided preliminary support for its use. The IRAD's Danger to Public and Unlikely to Appear domains showed good to excellent interrater reliability, and the latter domain predicted ATD violations with moderate accuracy. Concordance analyses revealed misalignments between client risk and CBSA's purportedly risk-based decisions. The IRAD and CBSA's current security classification tool also showed similar concordance with security classification decisions. Finally, we adapted the IRAD prototype for operational use, creating a 21-item structured professional judgement tool. We then explored the potential operational utility of this tool in a mixed prospective-retrospective pilot with CBSA employees. Though low officer engagement prevented robust evaluation, we found further evidence of misalignment between risk and decisions. A noise audit with client vignettes also revealed inconsistencies among officers during the risk identification, risk analysis, and decision-making processes for ATD determinations. As the IRAD consolidates CBSA's operational resources, our research suggests that decisions are influenced by other factors, which may be extraneous given the inconsistencies identified in our noise audit. The IRAD's streamlined, empirically informed design and preliminary evidential support may therefore help CBSA better align its decisions with risk.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.029 | 0.096 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.011 | 0.003 |
| Scholarly communication | 0.006 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".