Bibliographic record
Abstract
Federated Learning (FL) enables collaborative model training across multiple participants without sharing raw data, achieving a balance between privacy protection and data utility. However, the decentralized nature of FL also introduces new security threats, among which model poisoning attacks are the most critical. Malicious clients can upload manipulated model updates to disrupt the global aggregation process, leading to performance degradation or even system failure. This thesis systematically investigates anomaly detection (AD) and attack mitigation in FL, aiming to enhance system robustness and security from both detection and defense perspectives. First, a comprehensive review and categorization of AD methods are presented, focusing on reconstruction-based and prediction-based deep learning frameworks and their applications to different data types. Second, existing defense mechanisms are analyzed in terms of their effectiveness and limitations under various attack scenarios, providing the theoretical foundation for the proposed framework. Based on these analyses, this thesis proposes an unsupervised defense framework named Dual-VAE with Truncated Gaussian (DVTG). The framework follows a three-stage structure to model and filter client updates. In Stage 1, a variational autoencoder (VAE) is trained to estimate reconstruction errors and identify a set of potentially clean updates. In Stage 2, a second VAE with a truncated Gaussian prior is trained on this refined subset to obtain a more stable latent representation. In Stage 3, the trained model evaluates incoming client updates and filters those with high reconstruction errors before aggregation. The method enables effective anomaly detection without requiring labeled or clean data and remains stable under both adversarial and stochastic disturbances. Experiments conducted on the MNIST dataset under non-independent and identically distributed (non-IID) conditions show that DVTG outperforms the baseline model across different attack scenarios. The framework effectively detects malicious clients while maintaining stable convergence and comparable accuracy to the non-attack scenario. Finally, this thesis discusses several future directions. These include extending the defense framework to hierarchical FL architectures and developing more interpretable and efficient AD models. The goal is to build a reliable and practical FL framework with stronger defense capability and better adaptability to real-world environments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".