MétaCan
Menu
Back to cohort
Record W7053699420

Who Tests the Testers? Assessing the Effectiveness and Trustworthiness of Deep Learning Model Testing Techniques

2024· other· fr· W7053699420 on OpenAlexfundno aff

Bibliographic record

VenuePolyPublie (École Polytechnique de Montréal) · 2024
Typeother
Languagefr
FieldEngineering
TopicLaser Design and Applications
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of CanadaConsortium de Recherche et d’innovation en Aérospatiale au Québec
KeywordsTrustworthinessDeep learningContext (archaeology)Validation test
DOInot available

Abstract

fetched live from OpenAlex

RÉSUMÉ: L’arrivée des algorithmes d’apprentissage profond et leur intégration dans tous les domaines de la vie courante ont eu et continu d’avoir un impact important sur la société. Plus récemment, l’arrivée de l’Intelligence Artificielle générative à travers des technologies tel que Chat- GPT ont accéléré ce processus. En parallèle, plusieurs recherches ont montré les limitations de ces systèmes au regard de leurs possibles défaillances, posant la question de comment faire en sorte de prévenir ces problèmes pour améliorer leur fiabilité. Une manière établie de répondre à cette problématique consiste à tester ces systèmes, afin de détecter ces défaillances avant le déploiement de ces systèmes et trouver les fautes responsables et les réparer. À cet effet, plusieurs techniques de tests ont été développées au fil des années, en s’inspirant de techniques existantes dans le domaine du test logiciel ou en construisant de nouvelles méthodes adaptées à ces nouveaux algorithmes. Cependant, le développement de ces techniques a montré également que le nouveau paradigme apporté par l’apprentissage profond, i.e., le fait que la logique interne des modèles est “apprise” et non codée, a changé la donne. À travers cela, c’est la fiabilité et l’efficacité de ces méthodes de tests qui sont remises partiellement en question, ce qui compromet à son tour la fiabilité de ces algorithmes d’apprentissage profond et la confiance des utilisateurs. C’est dans ce contexte que se situe cette thèse. Ce travail, organisé en huit chapitres, vise à adresser la question de la fiabilité des techniques de tests appliqués aux modèles d’apprentissage profond via le développement de quatre cadriciels pour améliorer la fiabilité et l’efficacité de techniques de tests. Chacun d’entre eux est orienté sur un aspect différent en termes des limites des techniques de tests et du sous-paradigme d’apprentissage profond concernés, peignant différentes possibilités d’améliorer ces techniques de tests. Ce travail commence par une introduction (Chapitre 1) contextualisant le problème avant de définir les connaissances préalables nécessaires à cette thèse (Chapitre 2). Par la suite, une revue de la littérature illustre les différentes problématiques et techniques existantes liées à ce travail (Chapitre 3). Cela donne lieu à quatre chapitres décrivant chacun un cadriciel particulier. ABSTRACT: The rise of Deep Learning models and their integration into daily life have impacted society, and the recent trend of Generative Artificial Intelligence models has boosted this process. However, in parallel to those developments, several studies have shown the limitations of such models, notably the dramatic failures they incur, raising the question of their trustworthiness. One established way of dealing with this issue is to test those models so failures can be detected prior to deployment and the faults causing them to be identified and fixed. To that end, multiple testing techniques have been developed over the years, either adapted from traditional software testing or devised to adapt to those new models. However, the development of those techniques has shown that the new paradigm brought about by Deep Learning, that is, an inner logic “learned” and not coded, could impact the trustworthiness and effectiveness of the testing techniques itself, undermining the effort to foster users’ trust in Deep Learning models. This thesis takes place in that context and, organized in eight chapters, aims to deal with the trustworthiness of testing techniques applied to Deep Learning models by defining four frameworks for improving the trustworthiness and effectiveness of testing techniques in Deep Learning. Each framework is focused on a different aspect of the problem, both in terms of the limits of the testing techniques tackled and of the sub-paradigms of Deep Learning investigated, thus giving a comprehensive picture of the possible improvement. The thesis starts with an introduction (Chapter 1) to frame the problem the thesis is tackling and then provides prior knowledge of works dealt with in this work (Chapter 2). Then, a literature review (Chapter 3) is presented, describing related works and current problems related to the study. Finally, the following four chapters deal with the frameworks mentioned above. Chapter 4 starts by exploring the testing techniques targeting faults in the code and specifications of Deep Reinforcement Learning algorithms, extending the concept of mutation testing to this particular sub-paradigm of Deep Learning. This chapter illustrates the application issues of mutation technique in Deep Reinforcement Learning and shows a possible solution through the proposed framework RLMutation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.071
metaresearch head score (Gemma)0.369
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.929
Threshold uncertainty score0.378

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0710.369
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0010.002
Scholarly communication0.0030.004
Open science0.0020.003
Research integrity0.0030.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.011
GPT teacher head0.243
Teacher spread0.232 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuePolyPublie (École Polytechnique de Montréal)Same topicLaser Design and ApplicationsFrench-language works237,207