Who Tests the Testers? Assessing the Effectiveness and Trustworthiness of Deep Learning Model Testing Techniques
Bibliographic record
Abstract
RÉSUMÉ: L’arrivée des algorithmes d’apprentissage profond et leur intégration dans tous les domaines de la vie courante ont eu et continu d’avoir un impact important sur la société. Plus récemment, l’arrivée de l’Intelligence Artificielle générative à travers des technologies tel que Chat- GPT ont accéléré ce processus. En parallèle, plusieurs recherches ont montré les limitations de ces systèmes au regard de leurs possibles défaillances, posant la question de comment faire en sorte de prévenir ces problèmes pour améliorer leur fiabilité. Une manière établie de répondre à cette problématique consiste à tester ces systèmes, afin de détecter ces défaillances avant le déploiement de ces systèmes et trouver les fautes responsables et les réparer. À cet effet, plusieurs techniques de tests ont été développées au fil des années, en s’inspirant de techniques existantes dans le domaine du test logiciel ou en construisant de nouvelles méthodes adaptées à ces nouveaux algorithmes. Cependant, le développement de ces techniques a montré également que le nouveau paradigme apporté par l’apprentissage profond, i.e., le fait que la logique interne des modèles est “apprise” et non codée, a changé la donne. À travers cela, c’est la fiabilité et l’efficacité de ces méthodes de tests qui sont remises partiellement en question, ce qui compromet à son tour la fiabilité de ces algorithmes d’apprentissage profond et la confiance des utilisateurs. C’est dans ce contexte que se situe cette thèse. Ce travail, organisé en huit chapitres, vise à adresser la question de la fiabilité des techniques de tests appliqués aux modèles d’apprentissage profond via le développement de quatre cadriciels pour améliorer la fiabilité et l’efficacité de techniques de tests. Chacun d’entre eux est orienté sur un aspect différent en termes des limites des techniques de tests et du sous-paradigme d’apprentissage profond concernés, peignant différentes possibilités d’améliorer ces techniques de tests. Ce travail commence par une introduction (Chapitre 1) contextualisant le problème avant de définir les connaissances préalables nécessaires à cette thèse (Chapitre 2). Par la suite, une revue de la littérature illustre les différentes problématiques et techniques existantes liées à ce travail (Chapitre 3). Cela donne lieu à quatre chapitres décrivant chacun un cadriciel particulier. ABSTRACT: The rise of Deep Learning models and their integration into daily life have impacted society, and the recent trend of Generative Artificial Intelligence models has boosted this process. However, in parallel to those developments, several studies have shown the limitations of such models, notably the dramatic failures they incur, raising the question of their trustworthiness. One established way of dealing with this issue is to test those models so failures can be detected prior to deployment and the faults causing them to be identified and fixed. To that end, multiple testing techniques have been developed over the years, either adapted from traditional software testing or devised to adapt to those new models. However, the development of those techniques has shown that the new paradigm brought about by Deep Learning, that is, an inner logic “learned” and not coded, could impact the trustworthiness and effectiveness of the testing techniques itself, undermining the effort to foster users’ trust in Deep Learning models. This thesis takes place in that context and, organized in eight chapters, aims to deal with the trustworthiness of testing techniques applied to Deep Learning models by defining four frameworks for improving the trustworthiness and effectiveness of testing techniques in Deep Learning. Each framework is focused on a different aspect of the problem, both in terms of the limits of the testing techniques tackled and of the sub-paradigms of Deep Learning investigated, thus giving a comprehensive picture of the possible improvement. The thesis starts with an introduction (Chapter 1) to frame the problem the thesis is tackling and then provides prior knowledge of works dealt with in this work (Chapter 2). Then, a literature review (Chapter 3) is presented, describing related works and current problems related to the study. Finally, the following four chapters deal with the frameworks mentioned above. Chapter 4 starts by exploring the testing techniques targeting faults in the code and specifications of Deep Reinforcement Learning algorithms, extending the concept of mutation testing to this particular sub-paradigm of Deep Learning. This chapter illustrates the application issues of mutation technique in Deep Reinforcement Learning and shows a possible solution through the proposed framework RLMutation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.071 | 0.369 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".