Hardware Architectures for Lossless Compression
Bibliographic record
Abstract
Demands for storing huge volumes of data and limited communication networkbandwidth call for effective data compression with high performance and energy efficiency to\nreduce high storage and communication costs for a wide range of systems and applications.\nData compression consists of two types: lossy and lossless. Canonical Huffman encoding and\nGzip are two popular lossless compression techniques. Hardware implementation of such\nlossless compression techniques in many-core processors and other hardware platforms is\ncrucial in achieving optimized performance and energy-efficiency. This dissertation analyzes\nstatic and dynamic canonical Huffman codec algorithms, Gzip compression techniques, and\npresents high throughput and energy-efficient hardware and software implementations with\ngood compression ratios.Canonical Huffman encoding is naturally sequential which consists of several tasks:finding a histogram of symbol-frequencies, sorting of symbols, Huffman tree creation, building\ncanonical code tables, and encoding symbols. This dissertation demonstrates energy-efficient\ncanonical Huffman encoder architectures exploiting task-level parallelism for the above tasks\nand introduces a concurrent approach to execute sorting, Huffman tree creation, and code\nlength computation tasks that yields better memory efficiency and performance than the\nconventional approach. The proposed architectures are implemented on a many-core array,\nIntel i7, Nvidia GT 750M, Intel FPGA, and 45-nm ASIC.The many-core encoder implementations achieve a scaled throughput per chip areathat is 89.2x and 4.7x greater on average and 44.7x and 8.2x greater in terms of scaled\nenergy efficiency (compressed bits encoded per energy) than the Intel i7 and Nvidia GT\n750M, respectively executing the common Corpus benchmarks for data compression. Encoder\nimplementations on the many-core processor array yield scaled throughput per chip area and\nscaled energy efficiency that is 58x and 4.8x greater on average than the state-of-the-art\nefficient canonical Huffman encoder implementation on Tesla V100 GPUs executing the\nenwik8 dataset.Scaled synthesis results from a proposed pipelined canonical Huffman encoder 45-nm ASIC results in 2.44x greater on throughput, 3.5x lower on total power dissipation, and7.6x lower energy dissipation over Intel FPGA implementations.Next, this dissertation presents bit-parallel static and dynamic canonical Huffmandecoder implementations using an optimized lookup table approach on a fine-grain manycore\narray, Intel i7, Nvidia GT 750M, Intel FPGA, and 45-nm ASIC. The many-core\nimplementations achieve a scaled throughput per chip area that is 891x and 7x greater on\naverage and scaled energy efficiency (compressed bits decoded per energy) that is 149.5x and\n3.9x greater on average than the i7 and GT 750M, respectively.The 45-nm ASIC synthesis results show that the pipelined and memory-efficient staticdecoder yields a 5.1x throughput improvement and 13.4x energy efficiency improvement\nover the FPGA implementation.Furthermore, this dissertation presents energy-efficient and high throughput Gziphardware architectures using a chained hash bank memory design of depth three for LZ77\nencoder and synthesis results of the Gzip compression engines implemented in a 45-nm ASIC.\nThe proposed encoder architectures exploit both static and dynamic canonical Huffman\nencoders along with a pipelined LZ77 encoder. The pipelined Gzip engine using a static\ncanonical Huffman encoder with a parallel window size (PWS) of 16 bytes per clock cycle,\nachieves a maximum input throughput of 2.53 GB/s, while the dynamic canonical Huffman\nencoder-based Gzip compressor achieves a maximum input throughput of 0.52 GB/s. To the\nbest of our knowledge, this Gzip compressor offers the highest reported compression ratio in\nthe literature of 2.47 for the Calgary Corpus benchmark.Finally, this dissertation presents DeepScaleTool, an open-source tool for the accurateestimation of deep-submicron technology scaling by modeling and curve fitting published data\nby a leading commercial fabrication company for silicon fabrication technology generations\nfrom 130 nm to 7 nm for the key parameters of area, delay, and energy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".