MA2DF: A Multi-Agent Anomaly Detection Framework
Bibliographic record
Abstract
Time-sensitive safety-critical systems store traces as a collection of time-stamped messages that are generated while a system is operating. Analysis of these traces becomes a key task as it allows one to find faults or errors within a system that is otherwise difficult to discern, especially in complex systems. Furthermore, finding any form of anomalous behaviour becomes critical in time-sensitive and safety-critical systems where a late detection will often lead to dire consequences. Most available approaches are generally used in networking or business process analysis. We focus on creating a lightweight and explainable approach for time-sensitive safety-critical systems. By using a set of system traces under both normal and anomalous conditions, our approach attempts to classify whether or not a trace is anomalous. In this work, we introduce MA2DF, Multi-Agent Anomaly Detection Framework, a novel multi-agent based graph design approach for online and offline anomaly detection in system traces. Our approach takes advantage of the timing information between a sequence of events and also the event sequences to learn and discern between normal and anomalous traces. We present two approaches, an offline approach to discern anomalous behaviour by utilizing the event occurrence workflow graph. The second approach is an online streaming algorithm that monitors the sequence of events as they arrive in real-time. This can be used to detect anomalies, find the cause, and improve system resilience. We show how our approach, MA2DF, is superior to other state-of-the-art models. The paper will explore the technical feasibility and viability of MA2DF by utilizing industry strength case study using traces from a field-tested hexacopter.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".