Study 3 - non-verbal cues in sincere testimonies
Bibliographic record
Abstract
Prior research has extensively examined witnesses’ non-verbal cues in legal contexts, primarily focusing on the association between these cues and deception. Yet, findings reveal only weak and small effects between such cues and deception (e.g., DePaulo et al., 2003; Hartwig & Bond, 2011; Vrij & Granhag, 2012). Luke (2019) further cautions that these correlations might be overstated due to publication bias and methodological limitations, underscoring the challenges in distinguishing true effects from statistical artifacts in deception studies. Despite these insights, a notable gap remains in understanding how non-verbal cues relate to accuracy in honestly reported witness testimonies, an area that has received less attention. Honestly testimonies refer to recollections from memory without the intent to deceive. Currently, the impact of non-verbal cues on the accuracy and credibility of witness statements remains poorly understood. Krahmer and Swerts (2005) demonstrated that visual cues, such as changes in facial expressions, are indicative of a speaker's Feeling of Knowing (FOK), thus signaling potential uncertainty. These findings suggest that such cues could affect perceptions of a witness’ credibility in legal contexts, as they complement verbal communication and influence the observer's interpretation and understanding of the testimony (Esteve-Gibert & Guellaï, 2018). However, using non-verbal cues for credibility assessments in legal settings is fraught with challenges. For example, Denault et al. (2023) reveal that Canadian judges do rely on such cues—like eye contact, gestures, and tone—for credibility judgments, despite no empirical evidence of the reliability of these cues. This reliance, amidst the scientific uncertainties about interpreting these cues, underscores the problematic nature of their use in critical legal decisions. Taken together, observable markers that could differentiate accurate from inaccurate statements in honest testimonies are of both theoretical and practical interest, yet such distinctions have not been investigated before. This study builds on the findings of Raver et al. (2023), which revealed that non-native speaking witnesses reported lower confidence and were perceived as less credible than native speakers, even though their testimonies were equally accurate. This complexity suggests a detailed examination is needed. Therefore, we will set out to investigate if (H1) non-verbal cues are associated with the accuracy of statements, investigating whether display of such cues can reliably predict whether statements are correct or incorrect, and if so (H2) whether this differs between native vs. non-native speaking witnesses. Also, we will examine (H3) the relation between non-verbal cues and witnesses’ self-reported confidence, assessing how these cues correlate with the confidence witnesses report in their own statements. Lastly, (H4) we will investigate how the display of non-verbal cues relate to the perceived credibility of the witness by independent observers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.008 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.004 | 0.001 |
| Open science | 0.009 | 0.005 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.066 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".