Two Decades of Applied AI: A Research Journey Through Security, Networks, and Language
Bibliographic record
Abstract
This keynote traces a two-decade research journey at the intersection of machine learning, cybersecurity, human interaction, and natural language processing, shaped by the evolution of AI methods and their application to real-world challenges. While rooted in Dr. Vargas Martin’s research, the talk reflects more broadly on how AI has matured as a powerful enabler of scientific inquiry and system design across domains. The journey began in 2006 with statistical learning techniques for detecting child sexual abuse material in network traffic, an early demonstration of AI in support of digital safety. In 2011, neural networks were applied to predict learner behaviour in digital environments, paving the way for future user modeling efforts. By 2015, the focus shifted to detecting covert side-channel communication in wireless and mobile ad hoc networks, where machine learning uncovered hidden signaling patterns within low-level protocols. From 2018 to 2022, the research moved increasingly toward human-centered security, investigating password memorability and later generating resilient authentication data using adversarial learning and pre-trained language models. In parallel, new directions emerged in affective computing, including the modeling of artificial empathy in clinical companion robots with privacy-by-design principles (2021), and the development of emotion recognition systems for social robots (2022). Most recently, the work has returned to foundational NLP problems, including enhanced sentence-wise text segmentation using transformer models (2024) and the application of large language models to detect cryptographic misuse in software systems (2025). Throughout this arc, machine learning has remained a constant, not merely as a method, but as a lens through which to interpret, model, and shape intelligent, secure, and human-aware systems. This talk will explore that continuum, situating past projects within the evolving landscape of AI and drawing lessons for future interdisciplinary research in the spirit of the SNPD community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.006 | 0.027 |
| Scholarly communication | 0.016 | 0.025 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.007 | 0.019 |
| Insufficient payload (model declined to judge) | 0.011 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".