Neural Network Structural Pruning and Acceleration in Frequency Domain
Bibliographic record
Abstract
In computer vision, the expanding size of neural networks raises concerns about the potential information overload caused by the vast number of network parameters. Leveraging the fact that high-frequency image components are less critical, many parameters can be set to zero through frequency-domain unstructured pruning, which has proven effective in this context. While models with numerous zero parameters can be significantly compressed, reducing storage and transportation costs, the necessity of transforming inputs between frequency and spatial domain via discrete cosine transform (DCT) between layers imposes an additional computation. Despite the abundance of zero parameters, the total parameter count remains unchanged, and computational demands may even increase. To address this, we propose a novel method to dynamically prune the structure of frequency-domain models, achieving further compression and acceleration. Specifically, we exploit the linearity of the Fourier transform during conversion from the frequency domain to the spatial domain. By retaining only the low-frequency components of the model's parameters, we structurally prune channels corresponding to zeroed-out parameters. Evaluations on various network architectures, such as LeNet, AlexNet, VGG, ResNet, ViT, and UNet, confirm the efficacy of our structural pruning method. For instance, a ViT model with 47.86ms inference time has been accelerated to just 3.44ms, reducing theoretical computation by 85.70% and inference time by 92.82%, with an accuracy drop of only 0.78%.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".