Bibliographic record
Abstract
How do end users understand social engineering attacks, and how do their mental models differ from reality?To investigate, we have proposed a new social engineering attack framework, and ran two studies using the framework as the foundation.In the first study, we conducted 30 interviews to investigate social engineering mental models, and found that confidence and accuracy are underlying themes that affect users' mental models.In the second survey, we quantified how confidence and accuracy impact mental models at different stages of an attack.We found that users tend to be overconfident in their ability to understand social engineering attacks, but hold inaccurate beliefs.They hold major misconceptions of what constitutes as social engineering, and the threat levels of these attacks.Based on our results, we have proposed various educational and design opportunities to match social engineering mitigation strategies to end user mental models of social engineering.Screenshot of an SMS social engineering attack we showed participants.This example is a generalized attack trying to elicit fear to gain money from users. . . . . . . . . . . . . . . . . . .49 5.2 Screenshot of a social media social engineering attack we showed participants.This example is a targeted attack trying to appeal to greed to gain money. . . . . . . . . . . . . . . . . . . . . . .50 5.3 Accuracy rates by the stages of the framework. . . . . . . . . .55 5.4 Distributions of aggregated confidence and accuracy scores by the stages of the framework. . . . . . . . . . . . . . . . . . . .58 5.5 Distributions of aggregated confidence and accuracy scores by attack vector. . . . . .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".