A Class of Augmented Convolutional Networks Architectures for Efficient Visual Anomaly Detection
Bibliographic record
Abstract
Visual anomaly detection, the task of isolating visual data that do not conform to the defined notion of normality, is very crucial for the autonomous functioning of entities with exceptional potential in a spectrum of real-world applications. Prevalent methods of visual anomaly detection involve massive, complex, inefficient models whose performances are often restricted by the availability of data, the extent of hyper-parameter tuning and optimal model design. Moreover, popular deep learning approaches such as reconstruction-based methods that use a variant of AutoEncoders and generative methods like Generative Adversarial Network are not inherently designed for the task of anomaly detection. The above factors discussed raise the following severe problems: \n \n 1. The general model design may not be efficient without a dedicated anomaly detection objective hence lacking the ability to well distinguish anomalies from the normal data \n 2. The immense time and effort spent in the search of hyper-parameters and optimal model design restricts models to be immediately deployed for applications \n 3. The functioning of models involve a lot of human intervention and is data-centric preventing them to be used in automated, online detection tasks \n 4. The high performing, complex models are too huge to be used in edge applications with low computational capacity that require models with a low memory footprint \n \nTo overcome these issues, several modular, model-agnostic, efficient and novel improvements to conventional architectures have been proposed and suggested in this work and they can potentially be employed in any AutoEncoder based anomaly detection task. The focus of this work is to develop models that are simple, efficient, require low memory usage and reduced effort expended on hyperparameter tuning and the proposed improvements can aid in readily augmenting the performance over baseline models by a significant margin by producing robust, discriminative and discernible representations to help better segregate anomalies from normal samples. \n \nThe overall generic framework proposed throughout this research consists of multiple, efficient architectures that can be used for immediate deployment of models for practical, real-world automated anomaly detection tasks with minimal human intervention and to impart capabilities like online learning and self-regularization for best performance on image and video tasks. The superiority and efficacy of the proposed solutions are enunciated through quantitative and qualitative performance evaluation on a variety of image and video datasets from diverse domains along with rich visualization and ablation studies. This work also focuses on the exploration of interpretability in AutoEncoder-based anomaly detection models with modifications to adapt popular classifier-centric explainability frameworks, to pave way for a better understanding of the function and decision of the models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".