That's not a Good Idea: A Robot Changes Your Behavior Against Social Engineering
Bibliographic record
Abstract
Dangers in modern human society are commonly attributed to the safety of online activities. In the domain of cybersecurity, Social Engineering (SE) relates to how attackers manipulate and coerce their targets into divulging sensitive information. One major problem in designing social engineering defenses is making users aware they are being targeted. In the context of fostering human empowerment and building an inclusive society, we explore the possibility of leveraging social robot companions to provide improved protection for individuals and companies against cybersecurity attacks, specifically focusing on the realm of social engineering (SE) tactics. We asked participants to play an immersive interactive storytelling game, challenging them with risky and social-engineering-related decisions and monitoring their explicit (i.e., decisions) and implicit (i.e., mouse trajectories and facial expressions) behavior. After each decision, the Furhat tabletop robot intervened, always suggesting the not-selected option. We compared two Compliance Gaining Behaviors (CGBs) the robot could use, either leveraging affection with the participants or logical thinking. Overall, Furhat’s interventions increased the acceptance of risky and SE proposals. However, comparing the situations in which the robot tried to convince participants to avoid a social engineering request to those in which it tried to persuade them to accept it, the former was significantly more successful. Also, participants struggled with ignoring Furhat’s advice, as shown by their more uncertain mouse trajectories and negative emotional valence. From the latter results, we trained a Decision Tree model, based on mouse trajectory features only, to predict if participants would change their minds with an accuracy of 64.9%. Such defense mechanisms could help better understand users’ decision-making process in cybersecurity and social engineering, designing more helpful and supportive robot companions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".