Study of Radiation Effects on FPGA and GPU based Neural Networks Accelerator Designs
Notice bibliographique
Résumé
Due to the rapid development of semiconductor technology and the increasing complexity of integrated circuit (IC) designs, they have expanded into various fields. In particular, they are used in environments affected by radiation such as space exploration, medical devices, etc. Therefore, the reliability of IC devices in radiation-hazard environments has become an important issue. Energetic particles, such as protons, neutrons, and heavy ions, can penetrate the device and cause single-event effects (SEEs), resulting in transient voltage or current changes in the circuitry. These changes can cause data to be modified, generate errors in logic, or even cause system crashes. Various radiation-resistant techniques have been proposed in this area of research, such as triple modular redundancy (TMR) , error-correcting codes (ECCs), etc., but the effectiveness of these designs gradually decreases as the technology node continues to shrink. Therefore, as IC technology continues to evolve, improving its radiation resistance remains very valuable research.\n\nConvolutional neural networks (CNNs), as the most widely used model for deep learning, performs well in image recognition and target detection. Field Programmable Gate Arrays (FPGAs), with their high degree of parallelism, programmability, and low power consumption, are ideal platforms for efficient CNN computation, and the design of FPGA-based CNN accelerator has been widely studied and applied. However, the complexity and sensitivity to radiation effects of FPGAs make them less reliable in radiation environments, which in turn affects the inference accuracy of CNN models on FPGA accelerator. In this study, the radiation reliability of CNN models developed on FPGAs is comprehensively evaluated by using proton radiation, two-photon absorption (TPA) laser scanning, and software-level fault injection approaches. The most sensitive modules were firstly found by fully evaluating a LeNet-5 based FPGA accelerator using TPA laser scanning, and then adding a register-level selected TMR hardened design, achieving a 40\\% improvement in reliability while adding only 20\\% redundancy in utilizations, with experimental results validated by both laser and proton tests. \n\nIn addition, the reliability of CNN models under different architectural designs is compared. Two popular CNN accelerator architectures: streaming architecture (SA) and single computation engine (SCE), are implemented on our FPGA board. Experimental results show that SA-based CNNs require more hardware resources but exhibit superior resilience against single event upsets (SEUs). Without any Radiation Hardened by Design (RHBD) protection, SCE has an error rate approximately twice as high as SA. At the same time, the use of dynamic partial reconfiguration (DPR) method combined with soft-error mitigation (SEM) IP core was proposed (AutoDPR-SEM). It significantly improves the reliability of the model without increasing the inference timing of both model. This AutoDPR-SEM significantly improves CNN accelerators reliability, reducing the critical error rate by approximately 17.8 times in SCE and 14.8 times in SA. A software level simulation is also applied to validate the TPA experiment, showing similar trends of the testing results across all models. \n\nTransformer networks, as high-performing models in the field of natural language processing (NLP), have demonstrated excellent performance across various applications. With the increasing model size and computational demands, GPUs have become the main platform for accelerating the training and inference of transformer networks. GPUs are popular in the field of deep learning due to their powerful parallel processing capabilities and efficient utilization of computational resources. However, GPUs also face challenges from SEE in radiation environments. In this study, the reliability of the popular transformer model DistilBERT is first evaluated using the TPA laser platform. An innovative soft-error impact assessment scheme is proposed, comparing the Euclidean distance (L2 distance) generated by the tensor output of each layer when affected by soft errors. When the output tensor shows an L2 distance greater than 1.0 compared to the standard tensor unaffected by soft errors, the likelihood of generating incorrect classification results significantly increases. This is used as a criterion to introduce a selective temporal redundancy computation method, which is enabled only when the output of the layers of impact is larger than 1.0 L2 distance. This approach significantly improves the reliability of running DistilBERT on GPU platforms. Laser experimental results validate the effectiveness of this approach.\n\nIn summary, this research proposes and evaluates radiation-hardened designs for FPGA and GPU to enhance the reliability of neural networks in radiation-prone environments. Through these studies, the feasibility of achieving high-reliability computing in harsh environments is demonstrated, providing essential references for future radiation-hardened electronic system designs.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».