MétaCan
Menu
Back to cohort
Record W7066156550

High-throughput BERT Inference for ARM Big.LITTLE Multi-core Processors

2023· dissertation· en· W7066156550 on OpenAlexaff

Bibliographic record

VenueeScholarship@McGill (McGill) · 2023
Typedissertation
Languageen
FieldPhysics and Astronomy
TopicParticle Detector Development and Performance
Canadian institutionsMcGill University
Fundersnot available
KeywordsInferenceFeature (linguistics)Set (abstract data type)Key (lock)
DOInot available

Abstract

fetched live from OpenAlex

Transformer-based models such as Bidirectional Encoder Representations from Transformers (BERT) have achieved state-of-the-art accuracy in Natural Language Processing (NLP) tasks.Nevertheless, these models are extremely cumbersome and have low throughput in NLP inference.This is more challenging for on-edge inference due to edge devices' limited memory size and computational power.Therefore, we aim to improve the edge inference throughput of transformer-based models, which is critical for real-life applications that process multiple independent tasks concurrently on resource-constrained devices to provide a better user experience.In this work, we propose a heterogeneous pipelining framework called PipeBERT, built on TVM, for BERT models to utilize all available heterogeneous resources present in the ARM big.LITTLE CPU architecture, which is a common processing element in modern edge devices.PipeBERT splits BERT models into subgraphs, then maps these subgraphs across two heterogeneous ARM CPU clusters of cores (big and LITTLE).On the HiKey970 embedded platform and for BERT models, PipeBERT demonstrates on average 48.6% of Abstract ii higher inference throughput than running on four big cores, and an average 61% of lower Energy-delay Product (EDP) than the best homogeneous inference (that is running a BERT model on either big or LITTLE clusters).Moreover, for the first time, we incorporate Neural Architecture Search (NAS) with pipeline inference for DynaBERT models on HiKey970 board.DynaBERT is a reconfigurable embedded BERT model, which can flexibly adjust its models' size and latency.We demonstrate up to 9 higher inference throughput than running homogeneous inference on the BERT-base model with only 1.3% accuracy degradation when we incorporate NAS with pipeline on DynaBERT.However, incorporating NAS with the pipeline raises the question of whether to do (1) NAS-then-Pipeline or (2)Pipeline-then-NAS (apply pipeline inference and then perform NAS).NAS-then-Pipeline would help designers to save time since implementing and profiling hardware metrics for pipelined models in a large design space would be complex.However, we show that even though this convention saves time, it results in non-optimal design choices.The result shows that 30% of found Pareto-optimal Front (POF) sets in Pipeline-then-NAS are different from what is found in NAS-then-Pipeline in the throughput-accuracy design space.Moreover, Pipeline-then-NAS 's POF finds some solutions that have up to 34% higher throughput than the closest design points found in NAS-then-Pipeline POF set.This shows the necessity of applying pipeline inference before NAS for finding optimal throughput-accuracy solutions.iii AbrgLes modles bass sur les transformateurs tels que BERT ont atteint une prcision de pointe dans les tches de traitement du langage naturel (NLP).Nanmoins, ces modles sont extrmement encombrants et ont un faible rendement dans l'infrence NLP.Le dfi est encore plus grand pour les infrences on-edge en raison de la taille limite de la mmoire et de la puissance de calcul des appareils de bord.De ce fait, nous cherchons amliorer le rendement des infrence edge des modles bass sur les transformateurs, ce qui est essentiel pour les applications relles traitant simultanment plusieurs tches indpendantes sur des appareils aux ressources limites afin d'offrir une meilleure exprience aux utilisateurs.Dans ce travail, nous proposons un cadre de pipelining htrogne appel PipeBERT, construit sur TVM, pour que les modles BERT utilisent toutes les ressources htrognes disponibles prsentes dans l'architecture ARM big.LITTLE CPU, qui est un lment de traitement commun des appareils egde modernes.PipeBERT divise les modles BERT en sous-graphes, puis mappe ces sous-graphes sur deux groupes de coeurs de ARM CPU htrognes (big et LITTLE).Sur la plateforme intgre de HiKey970 et pour les modles List of Acronyms ADRS Average Distance From The Reference Set.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.840
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0010.000
Scholarly communication0.0000.001
Open science0.0010.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.043
GPT teacher head0.292
Teacher spread0.249 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueeScholarship@McGill (McGill)Same topicParticle Detector Development and PerformanceFrench-language works237,207