Reconfigurable Multi-Precision Multipliers for Energy- Efficient CNN Acceleration for Visual AI in ICT Systems
Bibliographic record
Abstract
Modern convolutional neural networks (CNNs) are essential in information and communication technology (ICT) applications, including edge computing, IoT devices, and mobile platforms, where energy efficiency and throughput are critical. These systems increasingly utilize multi-precision arithmetic to optimize accuracy and resource efficiency. However, traditional methods that assign separate fixed-precision multipliers for different bit-widths are inefficient, as the largest multiplier often dominates the critical path, limiting overall performance. In this paper, we introduce two scalable, power-efficient multiplier architectures with runtime reconfigurability: R4RC16 and R4RC32. These architectures are designed for CNN acceleration under multi-precision pruning. Each design features a low-power mode (8-bit) and a default mode (16-bit for R4RC16 and 32-bit for R4RC32), allowing for dynamic precision adjustment during inference with minimal overhead. When operating in low-power mode, our proposed multipliers achieve up to 7.6× greater energy efficiency compared to state-of-the-art approximate logarithmic multipliers, and up to 13.8× compared to approximate Booth-based designs. Additionally, they provide 2× (for 16-bit) and 4× (for 32-bit) higher throughput than exact 8-bit multipliers when processing pruned CNN workloads. Notably, the overhead in low-power mode is nearly independent of the full bit-width, resulting in a nearly constant power-delay product across both 16-bit and 32-bit designs. These findings highlight the significance of reconfigurable arithmetic units as critical components of ICT infrastructure that support healthcare, education, and multimedia, enabling CNNs to dynamically balance accuracy, energy, and throughput with less than 1% area overhead.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".