Power and Performance Based Autotuning of Heterogeneous Applications for CPU-GPU Systems
Bibliographic record
Abstract
<p>Recent advancement in artificial intelligence (AI) and deep learning has happened due to the usage of General Purpose Graphics Processing Units (GPUs) to implement these AI applications. GPU programming became easier with the advent of high-level abstraction API frameworks such as OpenCL and CUDA. The portability of these frameworks has been the performance cost. The GPU kernel performance is highly dependent on the underlying hardware architecture. The application kernels need their tuning every time it executes on a new device. The work presented in this thesis focuses on OpenCL kernels running on heterogeneous CPU-GPU systems. First, we present an analytical approach to estimate the power and performance of a convolution neural network (CNN) on a heterogeneous system that is useful for power and performance-based auto-tuning. Then we present our main contribution to multi-objective OpenCL kernels and propose an auto-tuner (MOKAT) for power and performance tuning. MOKAT tunes an OpenCL kernel without compromising on any of the two objectives (power and performance) and provides a final set of pareto-optimal kernels. MOKAT offers an integrated power calculation methodology for both online and offline tuning. It utilizes Non-Dominated Sorting Genetic Algorithm (NSGA-II) as the multi-objective evolutionary algorithm (MOEA). We describe the MOKAT API and internal structure of our framework. The two case studies of kernel tuning are related to 2D convolution and General Matrix Multiplication.</p>
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".