Towards multiplier-less implementation of neural networks
Notice bibliographique
Résumé
Artificial Intelligence (AI) has become profoundly embedded in contemporary life, with its applications proliferating across a wide array of domains.Central to AI are neural networks, which have markedly enhanced the capabilities of AI in areas such as computer vision and natural language processing.As neural networks scale in both size and computational complexity, the intelligent devices tasked with executing these networks face growing demands for computational and energy resources to ensure efficient and reliable performance.Consequently, resource-limited embedded devices, such as smartphones, encounter significant challenges in deploying state-of-the-art AI models.These devices frequently resort to cloud-based platforms, which necessitate continuous internet connectivity.However, cloud-based solutions pose several critical issues, including concerns over security, privacy, latency, and notably, their substantial environmental impact due to high energy consumption.This dissertation seeks to address these challenges by reducing the computational complexity of neural networks to facilitate their deployment on embedded devices.Specifically, it targets the primary source of computational burden and a major contributor to energy consumption in neural networks: high-precision multipliers (e.g., 16-bit or 8-bit multipliers).We propose novel implementations of neural networks that either markedly reduce the bitwidth of multipliers (to 4 bits or fewer) or entirely replace them with simpler logic operations (e.g., XNOR and shift operations).Moreover, reducing the bit-width of neural networks leads to decreased memory storage demands and reduced memory access, thereby contributing to a further reduction in overall energy consumption.In our initial implementation of neural networks, we present a novel approach for training multi-layer networks utilizing Finite State Machines (FSMs).In this approach, each FSM is interconnected with every FSM in both the preceding and subsequent layers.We demonstrate that the FSM-based network can effectively synthesize complex multi-input functions, such as 2D Gabor filters, and perform non-sequential tasks, such as image classification on stochastic streams, without the need for multiplications, given that FSMs are implemented solely through look-up tables.Building on i the FSMs' capability to handle binary streams, we propose an FSM-based model specifically designed for handling time series data, applicable to temporal tasks such as character-level language modeling.In our second implementation, we introduce an advanced stochastic computing (SC) representation termed the dynamic sign-magnitude (DSM) stream.This representation is specifically designed to enhance the precision of short-sequence SC-based multiplication.The DSM framework facilitates the substitution of conventional neural network multiplications with more efficient bitwise XNOR operations.By employing DSM, we achieve a substantial reduction in the required sequence length for SC-based neural networks, while maintaining accuracy levels comparable to existing methodologies.In our third implementation, we propose a new training framework for base-2 logarithmic quantization of neural networks.This framework quantizes weights into discrete power-of-two values by leveraging information about the network's weight distribution, specifically the standard deviation.This method allows us to replace computationally intensive high-precision multipliers with more efficient shift-add operations.Consequently, our quantized networks use approximately one-eighth the number of parameters compared to conventional high-precision networks, without compromising classification accuracy.Finally, in our latest implementation, we introduce a novel training framework that utilizes quantization techniques to facilitate the conversion between quantized networks and spiking neural networks (SNNs).SNNs are inherently devoid of multiplications, relying instead on addition and subtraction.This new framework offers an alternative approach for training SNNs.Specifically, we modify the SNN algorithm and mathematically demonstrate that after T time steps, the modified SNN approximates the behavior of a quantized network with T quantization intervals.This allows for the straightforward replacement of any SNN with its corresponding quantized network for training purposes.Given that the SNN and the quantized network share identical parameters, we can seamlessly transfer the parameters from the trained quantized network to the SNN without additional steps.I would also like to extend my sincere gratitude to the members of my supervisory committee, Professor Brett H. Meyer and Professor James J. Clark, for their invaluable guidance and constructive feedback throughout this journey.In particular, I am especially grateful to Professor Brett H. Meyer for his steadfast emotional support and encouragement, which motivated me during challenging times and helped me persevere.I wish to express my deepest appreciation to my brother
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,002 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,008 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».