Hardware Architectures for Lossless Compression
Notice bibliographique
Résumé
Demands for storing huge volumes of data and limited communication networkbandwidth call for effective data compression with high performance and energy efficiency to\nreduce high storage and communication costs for a wide range of systems and applications.\nData compression consists of two types: lossy and lossless. Canonical Huffman encoding and\nGzip are two popular lossless compression techniques. Hardware implementation of such\nlossless compression techniques in many-core processors and other hardware platforms is\ncrucial in achieving optimized performance and energy-efficiency. This dissertation analyzes\nstatic and dynamic canonical Huffman codec algorithms, Gzip compression techniques, and\npresents high throughput and energy-efficient hardware and software implementations with\ngood compression ratios.Canonical Huffman encoding is naturally sequential which consists of several tasks:finding a histogram of symbol-frequencies, sorting of symbols, Huffman tree creation, building\ncanonical code tables, and encoding symbols. This dissertation demonstrates energy-efficient\ncanonical Huffman encoder architectures exploiting task-level parallelism for the above tasks\nand introduces a concurrent approach to execute sorting, Huffman tree creation, and code\nlength computation tasks that yields better memory efficiency and performance than the\nconventional approach. The proposed architectures are implemented on a many-core array,\nIntel i7, Nvidia GT 750M, Intel FPGA, and 45-nm ASIC.The many-core encoder implementations achieve a scaled throughput per chip areathat is 89.2x and 4.7x greater on average and 44.7x and 8.2x greater in terms of scaled\nenergy efficiency (compressed bits encoded per energy) than the Intel i7 and Nvidia GT\n750M, respectively executing the common Corpus benchmarks for data compression. Encoder\nimplementations on the many-core processor array yield scaled throughput per chip area and\nscaled energy efficiency that is 58x and 4.8x greater on average than the state-of-the-art\nefficient canonical Huffman encoder implementation on Tesla V100 GPUs executing the\nenwik8 dataset.Scaled synthesis results from a proposed pipelined canonical Huffman encoder 45-nm ASIC results in 2.44x greater on throughput, 3.5x lower on total power dissipation, and7.6x lower energy dissipation over Intel FPGA implementations.Next, this dissertation presents bit-parallel static and dynamic canonical Huffmandecoder implementations using an optimized lookup table approach on a fine-grain manycore\narray, Intel i7, Nvidia GT 750M, Intel FPGA, and 45-nm ASIC. The many-core\nimplementations achieve a scaled throughput per chip area that is 891x and 7x greater on\naverage and scaled energy efficiency (compressed bits decoded per energy) that is 149.5x and\n3.9x greater on average than the i7 and GT 750M, respectively.The 45-nm ASIC synthesis results show that the pipelined and memory-efficient staticdecoder yields a 5.1x throughput improvement and 13.4x energy efficiency improvement\nover the FPGA implementation.Furthermore, this dissertation presents energy-efficient and high throughput Gziphardware architectures using a chained hash bank memory design of depth three for LZ77\nencoder and synthesis results of the Gzip compression engines implemented in a 45-nm ASIC.\nThe proposed encoder architectures exploit both static and dynamic canonical Huffman\nencoders along with a pipelined LZ77 encoder. The pipelined Gzip engine using a static\ncanonical Huffman encoder with a parallel window size (PWS) of 16 bytes per clock cycle,\nachieves a maximum input throughput of 2.53 GB/s, while the dynamic canonical Huffman\nencoder-based Gzip compressor achieves a maximum input throughput of 0.52 GB/s. To the\nbest of our knowledge, this Gzip compressor offers the highest reported compression ratio in\nthe literature of 2.47 for the Calgary Corpus benchmark.Finally, this dissertation presents DeepScaleTool, an open-source tool for the accurateestimation of deep-submicron technology scaling by modeling and curve fitting published data\nby a leading commercial fabrication company for silicon fabrication technology generations\nfrom 130 nm to 7 nm for the key parameters of area, delay, and energy.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,009 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».