Towards multiplier-less implementation of neural networks
Bibliographic record
Abstract
Artificial Intelligence (AI) has become profoundly embedded in contemporary life, with its applications proliferating across a wide array of domains.Central to AI are neural networks, which have markedly enhanced the capabilities of AI in areas such as computer vision and natural language processing.As neural networks scale in both size and computational complexity, the intelligent devices tasked with executing these networks face growing demands for computational and energy resources to ensure efficient and reliable performance.Consequently, resource-limited embedded devices, such as smartphones, encounter significant challenges in deploying state-of-the-art AI models.These devices frequently resort to cloud-based platforms, which necessitate continuous internet connectivity.However, cloud-based solutions pose several critical issues, including concerns over security, privacy, latency, and notably, their substantial environmental impact due to high energy consumption.This dissertation seeks to address these challenges by reducing the computational complexity of neural networks to facilitate their deployment on embedded devices.Specifically, it targets the primary source of computational burden and a major contributor to energy consumption in neural networks: high-precision multipliers (e.g., 16-bit or 8-bit multipliers).We propose novel implementations of neural networks that either markedly reduce the bitwidth of multipliers (to 4 bits or fewer) or entirely replace them with simpler logic operations (e.g., XNOR and shift operations).Moreover, reducing the bit-width of neural networks leads to decreased memory storage demands and reduced memory access, thereby contributing to a further reduction in overall energy consumption.In our initial implementation of neural networks, we present a novel approach for training multi-layer networks utilizing Finite State Machines (FSMs).In this approach, each FSM is interconnected with every FSM in both the preceding and subsequent layers.We demonstrate that the FSM-based network can effectively synthesize complex multi-input functions, such as 2D Gabor filters, and perform non-sequential tasks, such as image classification on stochastic streams, without the need for multiplications, given that FSMs are implemented solely through look-up tables.Building on i the FSMs' capability to handle binary streams, we propose an FSM-based model specifically designed for handling time series data, applicable to temporal tasks such as character-level language modeling.In our second implementation, we introduce an advanced stochastic computing (SC) representation termed the dynamic sign-magnitude (DSM) stream.This representation is specifically designed to enhance the precision of short-sequence SC-based multiplication.The DSM framework facilitates the substitution of conventional neural network multiplications with more efficient bitwise XNOR operations.By employing DSM, we achieve a substantial reduction in the required sequence length for SC-based neural networks, while maintaining accuracy levels comparable to existing methodologies.In our third implementation, we propose a new training framework for base-2 logarithmic quantization of neural networks.This framework quantizes weights into discrete power-of-two values by leveraging information about the network's weight distribution, specifically the standard deviation.This method allows us to replace computationally intensive high-precision multipliers with more efficient shift-add operations.Consequently, our quantized networks use approximately one-eighth the number of parameters compared to conventional high-precision networks, without compromising classification accuracy.Finally, in our latest implementation, we introduce a novel training framework that utilizes quantization techniques to facilitate the conversion between quantized networks and spiking neural networks (SNNs).SNNs are inherently devoid of multiplications, relying instead on addition and subtraction.This new framework offers an alternative approach for training SNNs.Specifically, we modify the SNN algorithm and mathematically demonstrate that after T time steps, the modified SNN approximates the behavior of a quantized network with T quantization intervals.This allows for the straightforward replacement of any SNN with its corresponding quantized network for training purposes.Given that the SNN and the quantized network share identical parameters, we can seamlessly transfer the parameters from the trained quantized network to the SNN without additional steps.I would also like to extend my sincere gratitude to the members of my supervisory committee, Professor Brett H. Meyer and Professor James J. Clark, for their invaluable guidance and constructive feedback throughout this journey.In particular, I am especially grateful to Professor Brett H. Meyer for his steadfast emotional support and encouragement, which motivated me during challenging times and helped me persevere.I wish to express my deepest appreciation to my brother
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".