Architecture and CAD Techniques for Efficient FPGA Implementation of Machine Learning and Other Applications
Notice bibliographique
Résumé
Field-programmable gate arrays (FPGAs) offer an alternative to application-specific integrated circuits (ASICs) that is attractive in scenarios where flexibility may be required or where chip volumes are not sufficiently high enough to justify the costs of a custom chip. The flexibility of FPGAs offer users the power/performance benefits of a custom hardware implementation, compared to software running on a processor, without committing to one specific design. However, the flexibility can lead to inefficiencies in terms of development time, area, performance, and power. FPGAs are used for a variety of applications and different optimizations can be applied to increase efficiency of FPGA implementations. This thesis considers techniques that can be applied to achieve efficient implementation of machine learning and other applications on an FPGA from the perspectives of architecture and computer-aided design (CAD). We consider the use of a high-level synthesis (HLS) tool to synthesize an accelerator for deep convolutional neural networks (CNNs) on an FPGA. We implement a complete end-to-end system running on an Arria 10 SoC FPGA. The accelerator implements zero-skipping and reduced-precision convolution with minimal impact on accuracy. We evaluate various versions of the accelerator through software changes and tool constraints alone. Then, we propose architecture changes to the carry-chain architecture in FPGAs to improve the resource utilization of binarized CNNs (BNNs). We add additional carry-chain circuitry that propagates sum instead of the carry. We demonstrate that we are able to reduce FPGA resource utilization while keeping the additional circuit area small. Lastly, we propose to re-map some of the look-up-tables (LUTs) to use the existing carry-chain architecture in order to increase performance by adding a post-LUT mapping step to the FPGA CAD flow. Using a subject graph that closely matches the underlying hardware, we are able to select critical paths to take advantage of the existing fast dedicated carry-chain routing.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».