MétaCan
Menu
Back to cohort
Record W2954641956 · doi:10.1002/cl2.68

PROTOCOL: Interview and Interrogation Methods and their Effects on Investigative Outcomes

2010· article· en· W2954641956 on OpenAlexaboutno aff
Christian A. Meissner, Allison D. Redlich, Sujeeta Bhatt, Susan Brandon

Bibliographic record

VenueCampbell Systematic Reviews · 2010
Typearticle
Languageen
FieldPsychology
TopicDeception detection and forensic psychology
Canadian institutionsnot available
FundersAmerican Psychological Association
KeywordsInterrogationProtocol (science)PsychologyMedicineAlternative medicinePolitical sciencePathology

Abstract

fetched live from OpenAlex

The request for a systematic review of the research on interviewing and interrogation methods is extremely timely and germane to current social events. Specifically, bright lights have been shone on both military and police investigation methods. The effectiveness of military interviewing, or human intelligence (HUMINT), has come under intense scrutiny because of the situations in Iraq and Afghanistan, and the heated debate over the use and efficacy of torture for educing intelligence (see Evans, Meissner, Brandon, Russano, & Kleinman, in press; Redlich, 2007). Just recently, information was released about the “enhanced interrogation” tactics used by the CIA with prisoners of war. At the same time, in the criminal justice arena, police interview and interrogation methods are being called into question because of the increased identification of false confessions and wrongful convictions. False confessions are an international problem that has been documented in almost every continent (see Kassin, Drizin, Grisso, Gudjonsson, Leo, & Redlich, 2010). In response, several countries, including the United Kingdom, Norway, New Zealand, and Australia, have changed interrogation practices from those that are guilt-presumptive to information-gathering in nature. The United States, Canada, and many Asian nations continue to utilize a guilt presumptive, accusatorial framework (Costanzo & Redlich, 2009; Leo, 2008; Ma, 2007; Smith, Stinson, & Patry, in press). The purpose of this systematic review is to evaluate information-gathering and interrogative (guilt-presumptive or accusatory) methods for persons suspected of committing crimes.1 One potential measure of effectiveness is diagnosticity. Interviewing methods can be considered “diagnostic” when they produce a higher ratio of true to false confessions and/or ability to detect accurate from inaccurate information. When assessing the effectiveness of questioning techniques on investigative outcomes, it is important to consider the accuracy of the outcome as well as the outcome itself. It is equally important to assess efficacy when suspects are both guilty and innocent (when known), as these two contexts may produce different levels of effectiveness. The information-gathering method of interviewing is typified by Great Britain's model. In 1984, because of a spate of high-profile false confessions, Great Britain enacted the Police and Criminal Evidence (PACE) Act of 1984 (Bull & Soukara, 2010; Home Office, 2003), which prohibited the use of psychologically manipulative techniques and mandated the recording of custodial interrogations. In 1993, the Royal Commission on Criminal Justice further reformed British interrogation methods by introducing the PEACE2 model. More specifically, the PEACE model focuses on developing rapport, explaining the allegation and the seriousness of the offense, emphasizing the importance of honesty and truth-gathering, and requesting the suspect's version of events. Suspects are permitted to explain the situation without interruption and questioners are encouraged to actively listen. This interview method has the goal of “fact finding” rather than that of obtaining a confession (with an emphasis on the use of open-ended questions), and investigators are expressly prohibited from deceiving suspects (Milne & Bull, 1999; Mortimer & Shepherd, 1999; Schollum, 2005). In part, the PEACE model is based on components of the Cognitive Interview (CI; Fisher & Geiselman, 1992). The CI was derived from basic memory research and involves a series of strategies and techniques. One of the principal techniques is context reinstatement (attempts to reinstate emotions, perceptions, and sequences of the event to-be-remembered). Another technique is to vary the order in which events are recounted. Often, the research that has been conducted on interviewing styles concentrates on the individual techniques/strategies and the theories underlying them. For example, Vrij, Mann, Fisher, Leal, Milne, and Bull (2008) recently tested whether recalling an event in reverse order (which, in theory, should be more difficult for liars than truth-tellers) influenced others' abilities to accurately detect deception. Although the effectiveness of the CI has been researched extensively, the majority of the research (but importantly, not all) and subsequent reviews of the research (e.g., Kohnken, Milnes, Memon, & Bull, 1999) have focused on witnesses and victims' reports of events, not suspects. The accusatorial method (as defined here) is typified by the U.S. model (Leo, 2008). It is generally contradictory to the information-gathering style in that it is confrontational and guilt-presumptive. In the U.S., police questioning of suspects consist of two phases. The first phase is the Interview phase (such as the “Behavioral Analysis Interview”, or BAI, see Inbau, Reid, Buckley & Jayne, 2001), in which the investigator is trained to conduct a non-accusatorial interview to determine whether the person of interest is indeed “the suspect” and should therefore be formally interrogated. A major part of this determination of guilt is a reliance on non-verbal behavioral cues and analyses of linguistic styles that are believed to indicate deception, but which consistently have been found by scientific methods to be unreliable (see Bond & DePaulo, 2006 for a review). Thus, by definition, U.S. interrogations are guilt-presumptive processes – they are focused upon extracting a confession by suspects who are believed to be guilty of the crime (Inbau et al., 2001; Meissner & Kassin, 2002, 2004). This second phase – the formal interrogation – consists of a variety of psychologically oriented, compliance gaining tactics. As summarized by Kassin and Gudjonsson (2004), interrogations involve: (a) custody and isolation, in which the suspect is detained in a small room and left to experience the anxiety, insecurity, and uncertainty associated with police interrogation; (b) confrontation, in which the suspect is presumed guilty and told (sometimes falsely) about the evidence against him/her, is warned of the consequences associated with his/her guilt, and is prevented from denying his/her involvement in the crime; and finally (c) minimization, in which a now sympathetic interrogator attempts to gain the suspect's trust, offers the suspect face-saving excuses or justifications for the crime, and implies more lenient consequences should the suspect provide a confession. One important and particularly controversial difference between information-gathering and accusatorial methods is the permissible use of trickery and deceit (i.e., lying to suspects about evidence). The scientific study of investigative interviewing has proliferated in the past two decades or so. Both the PACE and PEACE models and some of their individual components (e.g., strategic disclosure of evidence, use of open-ended questions) have been studied in the field and in the laboratory (Bull & Soukara, 2009; Meissner, Russano, & Narchet, 2010; see also Clarke & Milne, 2001 for an evaluation). Similarly, numerous experiments have been conducted on general (e.g., minimization and maximization; Russano, Meissner, Narchet, & Kassin, 2005) and more specific accusatorial methods (e.g., presenting false evidence; Redlich & Goodman, 2003). However, to our knowledge, a synthesized review such as the one proposed here has not been undertaken. That is, a review focusing on the effectiveness of information-gathering and accusatorial methods of questioning suspects has yet to be done, but one that will surely be instructive to academics and policymakers alike. In the U.S., police and military interrogation and intelligence gathering methods, which are accusatorial and guilt-presumptive in nature, are under fire. Some 20 years ago, Great Britain underwent similar controversies and in response, made sweeping policy changes that arguably preceded the scientific research. Via this systematic review of the experimental literature and its subsequent results, we are in a unique position to inform public policy before it may be altered. As detailed below, our main methodology of review will be meta-analysis and study space analysis, which are appropriate when aiming to translate research into policy recommendations. The objective of this review is to systematically and comprehensively review published and non-published, experimental and quasi-experimental studies on the effectiveness of interviewing and interrogation methods. We plan to focus on suspects as our population, interview style (information-gathering, accusatorial, control) as the intervention, and the diagnosticity of the methods as the primary measure of efficacy. Our guiding question is whether information-gathering or accusatorial methods are more diagnostic in the accuracy of the information that is produced when employed on guilty and innocent suspects. When relevant and available, completeness and consistency of information will also be examined as indicators of effectiveness. Finally, important knowledge has been gained from field, quasi-experimental studies. Thus, although accuracy of outcomes (e.g., confessions) cannot be discerned, we conduct a separate systematic review of field studies. With interrogation and intelligence gathering methods under intense scrutiny, jurisdictions, states, and countries may have to revisit their questioning procedures and policies, if they have not done so already. As mentioned above, numerous nations in the recent past have changed their interrogation practices. Armed with a review such as the one proposed here, policy makers and law enforcement decision-makers will have the best information available. Interview methods that increase the amount of true information gained from actual perpetrators while at the same time do not increase false information from innocent individuals are important to identify via a systematic, comprehensive, and scientific approach. We will conduct two separate meta-analyses and a study space analysis. As we describe more fully below, a study space analysis is one that highlights the topics—and intersection of topics—that have and have not yet been (but need to be) studied. The primary products of this systematic review will be two meta-analyses (MA) and subsequent forest plots that graph the magnitude and direction of calculated effect sizes. The first MA (i.e., MA 1) will only include experimental studies in which the “ground truth” (i.e., whether the person is innocent or guilty) is known. The second MA (i.e., MA 2) will include quasi-experimental, field studies in which the ground truth is unknown. A study that may be eligible for MA 1 and the SSA would be one conducted by Vrij and colleagues (2007; see below for more details). A study that would not be eligible would be one by Colwell, Hiscock-Anisman, Memon, Rachel, and Colwell (2007) because it focuses on witnesses rather than suspects. In searching the below references and databases, we will determine relevance by reading titles and abstracts. For example, titles that clearly refer to victim/witness accounts will not be included. When more information is needed, we will access and review full reports. Graduate students will be responsible for determinations of initial relevance, with Professors Meissner and Redlich making final decisions. We will search for published and unpublished, experimental and quasi-experimental studies on information-gathering-accusatory interviewing. Although we will not set an a priori limit on the publication dates of searched studies, we anticipate that the majority of studies that would be potentially eligible (pass the first round of review) will have been published from 1980 to the present. We will use the following keywords to initiate the search. We expect more keywords to be generated as the search progresses. In addition, we will combine keywords to produce more targeted searches, such as “interview and suspect,” and “confession and interrogation.” Finally, the reviewers have many well-established contacts with researchers studying interviewing and interrogation here in the U.S. and abroad. In Appendix A, we have started a list of possible researchers to contact. We will reach out to known and unknown contacts for unpublished or ‘in press’ studies to possibly include. We have obtained the programs of the 2nd and 3rd (June 2008) International Conferences on Investigative Interviewing. Included in these programs are more than 100 presentations that we can follow up on to determine if the studies have been written up. First, we describe an experimental laboratory study that may be eligible for inclusion in MA 1. Then, we describe an example of a quasi-experimental field study that may be appropriate for MA 2. The typical experimental research paradigm on the interviewing of suspects involves first, a mock crime, and second, an interview session. Participants are usually randomly assigned to be guilty or innocent of the crime (or to tell the truth or lie), and then randomly assigned to one of two or more interview styles (or specific interview techniques). Experiments are usually recorded and outcomes are reliably coded. We use a recent study by Vrij and colleagues (2007) to illustrate. The title of the study was, Cues to deception and ability to detect lies as a function of police interview styles, and published in a leading journal, Law and Human Behavior. In Experiment 1, 120 college students participated; half of them participated in a staged event (playing the game Connect 4 with a confederate) in which money was taken from the wallet of another confederate. This was the “truth tellers” condition. In the other condition, the “liars” did not partake in this staged event, but instead were given scripted information about the event. The “liars” also were instructed to take the money out of the wallet, hide it on themselves, and pretend to have participated in the staged event. Next, both liars and truth tellers were told they would be interviewed and to convince the interviewer that they did not take the money. There were three interview conditions; 60 truth tellers and 60 liars were randomly assigned across them. In the “information-gathering” condition, participants were instructed to tell everything they could about the Connect 4 game, providing as much detail as possible, and follow-up questions were open-ended (as opposed to leading). In the “accusation” condition, participants were asked 11 questions with an accusatory tone, such as “Are you sure you're telling me the truth?” and “Your reactions make me think you're hiding something from me.” In the “behavior analysis interview” condition, participants were asked for free recall and then asked 15 BAI questions, such as “Do you think that someone else did purposefully take the money?” The interviews were then coded and scored using Criterion Based Content Analysis (CBCA) and Reality Monitoring (RM). These scores were used as the dependent measures. In brief, they found that accusatory interviews resulted in no discernible differences between truth tellers and liars when either CBCA or RM scores were examined, whereas the other two interviewing styles did (though not across the board). In Experiment 2, Vrij et al. (2007) showed the videotaped stimuli (the three interview conditions) to 68 British police officers. The officers were told that they would see clips of interviews of students who were lying or telling the truth. Officers made dichotomous judgments of accuracy, which were used to calculate hits and false positives. In brief, they found that accuracy was unaffected by interview style and that a truth bias was found in that truth telling was more accurately assessed than lying. Effect sizes were reported. Additional studies have followed similar experimental methods, such as Hartwig et al. (2005) and Vrij et al. (2008 and 2009). An example of a quasi-experimental study would be Study 2 by Bull and Soukara (2010). In this study, the authors coded 80 actual interviews of suspects, which were randomly selected from a sample of 200 interviews. The authors reliably coded the interviews for the presence/absence of 17 tactics, the extent to which suspects moved towards confession (1 = no change to 5 = move from denial to confession), and a dichotomous confession outcome. The 17 tactics were categorized into information-gathering (e.g., open questions, gentle prods) or interrogatory (e.g., maximization, intimidation, leading questions). As stated by the authors, “this Study 2 did not find a simple relationship between of and extent of to However, Study 2 did find that the two tactics of and leading questions interrogatory were more in the interviews These to the techniques that produce outcomes the accuracy of outcomes are with actual crime suspects. eligible studies have outcomes, we will an effect for outcome measure studies that a or information-gathering a effect will be calculated across these or a will be selected for inclusion in the At this (i.e., to our our of the literature is that if not studies have not interviewed participants over Thus, we do not anticipate outcomes from time to be However, if we do we will use the final outcome to effect sizes. the as above, a study will be This will include basic information about the study (such as and and and a for the of eligible studies is they will be and assigned a unique study basic information about the of publication and the study of will be we will also assess the of the studies. For example, we could use the by and recently by and colleagues (see This studies into of At a we will and dependent for are some of to be coded will on the of studies that and on them. We will two meta-analyses and a study space analysis. For both we will our outcome and assess the of effect sizes for outcome using a including a of the effect and the A analysis of will also be if a of studies is (with a given set of A forest the calculated effect sizes will be We will also to conduct a study space analysis, which is an analysis that a and of the current literature the relationship between dependent and these are and when the study space analysis is the which have been or are in need of study Thus, whereas literature reviews to focus on the of scientific studies, study space reviews can is to be which can be important for An example of a study space et al., 2008; 1) is As by and colleagues are in a study space 1) identify the studies is also part of the 2) for study a which the and dependent identify and into the a in the to an intersection of study for and the individual study into one et al. (2008) conducted a study space analysis for identification and for the of on The is below as an example of a studies will not be eligible for the study space and meta-analyses because they do not the experimental or quasi-experimental However, when such studies are in our we will make determinations about their possible relevance in outcome in developing research questions, and in the of These determinations will be made by and the studies for A of studies in our search will be in the final The below of dates will be to as as The review will be every three to The reviewers and their students will be responsible for the same search and methods will be of the reviewers have of Meissner has the deception and including several meta-analyses in this and other & Meissner & Kassin, Meissner, & 2008; & Meissner, 2005). has also a by the on investigative interviewing. This into a and and policy which was published by the & Meissner, 2010). Redlich has the literature on U.S. police and military as part of the scientific review to a on police interrogations and false confessions (see Kassin et al., 2010). This review (as well as Redlich, 2007; Redlich & Meissner, provide the for interest in to conduct this systematic for the CI & U.S. and CI & U.S. will also on the as Both were in a recent review of the U.S. Brandon, & Kleinman, 2009). and have knowledge and access to reports of by the U.S. (as well as the United relevant to this proposed will also in the reviewers to be in at the of at 1 from et al., study space & from et al., study space for and &

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.848
Threshold uncertainty score0.803

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.133
GPT teacher head0.462
Teacher spread0.329 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2010
Admission routes1
Has abstractyes

Explore more

Same venueCampbell Systematic ReviewsSame topicDeception detection and forensic psychologyFrench-language works237,207