Tiny Federated Wireless Foundation Models for Resource-Constrained Devices
Bibliographic record
Abstract
Deploying large-scale foundation models (FMs) in resource-constrained devices presents critical challenges due to their substantial computational and memory requirements. This is particularly relevant for multi-task wireless sensing FMs running on sensors. To overcome these limitations, we propose a tiny federated wireless foundation model (WFM) framework that combines spectrogram-guided structured block-wise pruning with federated learning (FL) for efficient on-device deployment. Our approach prunes non-essential encoder blocks in vision transformers (ViTs) by leveraging the masked spectrogram modeling (MSM) pretraining loss as an importance indicator, ensuring only the most structurally significant components are retained. This enables federated adaptation with frozen backbones and lightweight, task-specific heads, minimizing both computational burden and communication overhead. The pruning strategy preserves the integrity of spectrogram reconstruction, while federated fine-tuning supports decentralized learning across clients with heterogeneous data distributions. Experimental results on human activity sensing and radio signal identification tasks confirm the efficacy of our approach. Specifically, the pruned ViT-based WFMs achieve up to 93% multiply-accumulate operations (MACs) reduction, 85% lower CPU inference time, and 49% reduction in communication overhead, all while maintaining high task accuracy. Our method demonstrates strong generalization and robustness across varying pruning ratios and data heterogeneity levels, while substantially reducing communication overhead, making it highly suitable for real-world industrial IoT deployments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".