Improving On-Device Artificial Intelligence by Deploying Efficient, Robust, and Continuous Models
Bibliographic record
Abstract
Artificial intelligence (AI) has emerged as a revolutionary technology in recent years, transforming various industries and fields.AI technologies have traditionally relied on cloud computing, but with the explosive growth of digital devices, on-device AI has become an essential tool for providing more powerful and convenient services than ever before.On-device AI plays an important role in applications because it can provide low latency, better privacy and more personalized experiences for users.However, deploying AI technologies, particularly deep neural networks (DNNs), on common devices still faces several challenges.One of the main challenges is the limited computing resources available on devices.The computations required for AI technologies are complex and intensive, while the devices have limited computing resources that are insufficient to support the operations of AI technologies.Another significant challenge is the limited availability of data on devices.DNNs require large amounts of labelled data to be trained, and it can be difficult and expensive to gather sufficient data on devices, resulting in less accuracy and robustness of DNNs.ii In this thesis, we explore the aforementioned challenges associated with deploying AI technologies, particularly DNNs, on devices.Specifically, we aim to improve the reliability of DNNs working on devices by deploying efficient, robust, and continuous models.In the first thread, we focus on reducing the inference cost of DNNs through pruning unimportant parameters of models.We analyze the statistical information of feature maps in convolutional neural networks and propose deleting feature maps that have low information diversity and high similarities.In the second thread, we aim to enhance the robustness of DNNs against distribution (domain) shifts in data under two scenarios.In the scenario where the domain label is available, we propose a novel domain adversarial adaptation method that combines various fine-grained domain alignment techniques.These alignments include aligning both marginal and conditional distributions of features in different domains, concentrating the feature towards the center for each class in the target domain, and forcing every target domain class to better align with its source domain counterpart.In the scenario where the domain label is unavailable, we develop a new method that integrates two major modules.One module is a data augmenter that introduces data diversity and enhances the utilization of labelled data.The other module is a joint classification-reconstruction structure that trains a domain-adaptive classifier to adjust itself to newly collected unlabelled data.In the third thread, we strive to make DNNs work continuously without forgetting previously learned knowledge when encountering streaming data.To achieve this, we define the important strength (used to consolidate the previously learned knowledge) as the sensitivity of the iii global loss function to the model parameters and propose adjusting the important strength adaptively to align it with the dynamic parameter updates.We evaluated the effectiveness of the methods mentioned above on multiple publicly available datasets and benchmarks.Our results show that our proposed methods outperform the current state-of-the-art techniques, highlighting the potential of our approaches.iv Abrégé L'intelligence artificielle (IA) est devenue une technologie révolutionnaire ces dernières années, transformant diverses industries et domaines.Les technologies d'IA ont traditionnellement dépendu de l'informatique en nuage, mais avec la croissance explosive des appareils numériques, l'IA sur appareil est devenue un outil essentiel pour offrir des services plus puissants et plus pratiques que jamais auparavant.L'IA sur appareil joue un rôle important dans les applications car elle peut offrir une faible latence, une meilleure confidentialité et des expériences plus personnalisées pour les utilisateurs.Cependant, le déploiement des technologies d'IA, en particulier des réseaux de neurones profonds (DNN), sur des appareils courants reste confronté à plusieurs défis.L'un des principaux défis est la disponibilité limitée de ressources informatiques sur les appareils.Les calculs requis pour les technologies d'IA sont complexes et intensifs, tandis que les appareils disposent de ressources informatiques limitées qui sont insuffisantes pour prendre en charge les opérations des technologies d'IA.Un autre défi important est la disponibilité limitée de données sur les appareils.Les DNN nécessitent de grandes quantités de données xviii
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".