MétaCan
Menu
Back to cohort
Record W4210455774 · doi:10.1145/3490489

NPC: <u>N</u> euron <u>P</u> ath <u>C</u> overage via Characterizing Decision Logic of Deep Neural Networks

2022· article· en· W4210455774 on OpenAlexfundno aff
Xiaofei Xie, Tianlin Li, Jian Wang, Lei Ma, Qing Guo, Felix Juefei-Xu, Yang Liu

Bibliographic record

VenueACM Transactions on Software Engineering and Methodology · 2022
Typearticle
Languageen
FieldComputer Science
TopicAdversarial Robustness in Machine Learning
Canadian institutionsnot available
FundersJapan Society for the Promotion of ScienceNatural Sciences and Engineering Research Council of CanadaJST-Mirai ProgramNational Satellite of Excellence in Trustworthy Software Systems, National University of SingaporeNational Research FoundationBộ Giáo dục và Ðào tạoMinistry of Education - SingaporeNational Research Foundation SingaporeCanadian Institute for Advanced Research
KeywordsComputer scienceArtificial intelligencePath (computing)Machine learningArtificial neural networkGraphDeep neural networksDecision treeMirroringTheoretical computer science

Abstract

fetched live from OpenAlex

Deep learning has recently been widely applied to many applications across different domains, e.g., image classification and audio recognition. However, the quality of Deep Neural Networks (DNNs) still raises concerns in the practical operational environment, which calls for systematic testing, especially in safety-critical scenarios. Inspired by software testing, a number of structural coverage criteria are designed and proposed to measure the test adequacy of DNNs. However, due to the blackbox nature of DNN, the existing structural coverage criteria are difficult to interpret, making it hard to understand the underlying principles of these criteria. The relationship between the structural coverage and the decision logic of DNNs is unknown. Moreover, recent studies have further revealed the non-existence of correlation between the structural coverage and DNN defect detection, which further posts concerns on what a suitable DNN testing criterion should be. In this article, we propose the interpretable coverage criteria through constructing the decision structure of a DNN. Mirroring the control flow graph of the traditional program, we first extract a decision graph from a DNN based on its interpretation, where a path of the decision graph represents a decision logic of the DNN. Based on the control flow and data flow of the decision graph, we propose two variants of path coverage to measure the adequacy of the test cases in exercising the decision logic. The higher the path coverage, the more diverse decision logic the DNN is expected to be explored. Our large-scale evaluation results demonstrate that: The path in the decision graph is effective in characterizing the decision of the DNN, and the proposed coverage criteria are also sensitive with errors, including natural errors and adversarial examples, and strongly correlate with the output impartiality.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.007
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.009
Threshold uncertainty score0.017

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.007
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.001
Science and technology studies0.0000.002
Scholarly communication0.0020.002
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.040
GPT teacher head0.287
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations57
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueACM Transactions on Software Engineering and MethodologySame topicAdversarial Robustness in Machine LearningFrench-language works237,207