Evaluating Generative Reasoning Models for Credential Tweaking and Lightweight Client-Side Defense in IoT Ecosystems
Bibliographic record
Abstract
Generative reasoning models introduce a new paradigm in cybersecurity, enabling not only novel defenses but also sophisticated attack simulations. This paper investigates the use of open-source reasoning models to simulate credential tweaking behavior and enhance password-based authentication security in IoT environments. We propose Hybrid Similarity Scoring (HSS) and its user-contextualized variant HSSuser, a lightweight, client-side similarity metric combining structural (Damerau-Levenshtein) and character-distribution (cosine similarity) components to detect password reuse and subtle modifications or tweaks in real time. Following NIST guidelines, we analyzed over 4 billion password pairs from breached datasets and used five prompt designs in various reasoning models such as DeepSeek-R1, Qwen-QwQ, Phi4-Reasoning, Qwen3, and Magistral series to generate password variants mimicking attacker strategies. Experimental results show that reasoning models can produce highly similar modifications resembling real-world password reuse patterns, while prompt reframing significantly reduces risky outputs. HSS effectively quantifies these behaviors and is suitable for deployment in constrained IoT devices, offering an intent-aware, proactive layer of client-side defense against AIenhanced credential attacks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".