High-throughput BERT Inference for ARM Big.LITTLE Multi-core Processors
Bibliographic record
Abstract
Transformer-based models such as Bidirectional Encoder Representations from Transformers (BERT) have achieved state-of-the-art accuracy in Natural Language Processing (NLP) tasks.Nevertheless, these models are extremely cumbersome and have low throughput in NLP inference.This is more challenging for on-edge inference due to edge devices' limited memory size and computational power.Therefore, we aim to improve the edge inference throughput of transformer-based models, which is critical for real-life applications that process multiple independent tasks concurrently on resource-constrained devices to provide a better user experience.In this work, we propose a heterogeneous pipelining framework called PipeBERT, built on TVM, for BERT models to utilize all available heterogeneous resources present in the ARM big.LITTLE CPU architecture, which is a common processing element in modern edge devices.PipeBERT splits BERT models into subgraphs, then maps these subgraphs across two heterogeneous ARM CPU clusters of cores (big and LITTLE).On the HiKey970 embedded platform and for BERT models, PipeBERT demonstrates on average 48.6% of Abstract ii higher inference throughput than running on four big cores, and an average 61% of lower Energy-delay Product (EDP) than the best homogeneous inference (that is running a BERT model on either big or LITTLE clusters).Moreover, for the first time, we incorporate Neural Architecture Search (NAS) with pipeline inference for DynaBERT models on HiKey970 board.DynaBERT is a reconfigurable embedded BERT model, which can flexibly adjust its models' size and latency.We demonstrate up to 9 higher inference throughput than running homogeneous inference on the BERT-base model with only 1.3% accuracy degradation when we incorporate NAS with pipeline on DynaBERT.However, incorporating NAS with the pipeline raises the question of whether to do (1) NAS-then-Pipeline or (2)Pipeline-then-NAS (apply pipeline inference and then perform NAS).NAS-then-Pipeline would help designers to save time since implementing and profiling hardware metrics for pipelined models in a large design space would be complex.However, we show that even though this convention saves time, it results in non-optimal design choices.The result shows that 30% of found Pareto-optimal Front (POF) sets in Pipeline-then-NAS are different from what is found in NAS-then-Pipeline in the throughput-accuracy design space.Moreover, Pipeline-then-NAS 's POF finds some solutions that have up to 34% higher throughput than the closest design points found in NAS-then-Pipeline POF set.This shows the necessity of applying pipeline inference before NAS for finding optimal throughput-accuracy solutions.iii AbrgLes modles bass sur les transformateurs tels que BERT ont atteint une prcision de pointe dans les tches de traitement du langage naturel (NLP).Nanmoins, ces modles sont extrmement encombrants et ont un faible rendement dans l'infrence NLP.Le dfi est encore plus grand pour les infrences on-edge en raison de la taille limite de la mmoire et de la puissance de calcul des appareils de bord.De ce fait, nous cherchons amliorer le rendement des infrence edge des modles bass sur les transformateurs, ce qui est essentiel pour les applications relles traitant simultanment plusieurs tches indpendantes sur des appareils aux ressources limites afin d'offrir une meilleure exprience aux utilisateurs.Dans ce travail, nous proposons un cadre de pipelining htrogne appel PipeBERT, construit sur TVM, pour que les modles BERT utilisent toutes les ressources htrognes disponibles prsentes dans l'architecture ARM big.LITTLE CPU, qui est un lment de traitement commun des appareils egde modernes.PipeBERT divise les modles BERT en sous-graphes, puis mappe ces sous-graphes sur deux groupes de coeurs de ARM CPU htrognes (big et LITTLE).Sur la plateforme intgre de HiKey970 et pour les modles List of Acronyms ADRS Average Distance From The Reference Set.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".