Machine learning based thermal analysis of on-chip and chiplet-based systems
Bibliographic record
Abstract
Following Moore's Law, the number of transistors on chips has continued to increase, Dennard Scaling however has not kept pace.The increase in power density caused by this phenomena has lead to increased temperatures which must be addressed, as high temperatures lead to unfavorable effects on system performance and reliability.A key challenge in the design process is therefore to identify problematic designs early, to avoid significant computation and time spent on designs which are thermally unviable.However, traditional thermal analysis tools such as finite element method (FEM) solvers or compact thermal models (CTM), are computationally costly and time-consuming thus proving unsuitable for iterative processes, such as those at early stages in the chip design process.Researchers, therefore, have been striving to develop fast and accurate methods of predicting the chip temperature.Recently machine learning (ML) based solutions have shown great promise in a variety of electronic design automation (EDA) applications, including that of design space reduction and exploration.Neural networks (NN) have proven to be especially effective, due to their ability to accurately and efficiently learn and Abstract ii embed the relationships between design parameters and performance metrics.When applied for thermal analysis tasks these models are able to predict at much faster rates then FEMs and CTMs making them more suitable for iterative processes, such as those found at early stages in the chip design process.For the thermal analysis task, specific types of NNs are used, typically either convolutional neural networks (CNN) or graph neural network (GNN) based architectures.These types of models are ideal due to their ability to learn based off the locality of elements, a parameter that is especially important in thermal phenomena.To accurately predict the temperature, proper data structures are required.Common implementations utilize data from early stages, such as functional block placements and power density maps to predict hot spots or thermal maps.This lack of sophisticated data structures coupled with the absence of training datasets that reflect realistic system-on-chip and chiplet-based designs limits the applicability and generalizability of these models.In this thesis, the effects of thermal properties and thermal design power on a systems overall performance are discussed.Then, there is a more focused discussion on thermal analysis of on-chip and chiplet-based systems and related literature.Afterwards, thermally aware design processes are examined, followed by the application of machine learning models in EDA, and specifically in thermal analysis.The following chapter focuses on benchmark datasets and datasets used in the training of these models, and a synthetic dataset is introduced that allows for the training of generalizable machine learning models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".