HYRIDE: HYbrid and Robust Intrusion DEtection approach for enhancing cybersecurity in Industry 4.0
Bibliographic record
Abstract
The interconnectedness and smartness aspect between several components of Industry 4.0 has caused sudden increase in data and its exchange, which has resulted in significant cybersecurity challenges. Thus, a better threat intelligence technique is required for monitoring and identifying malicious cyberattacks. However, distinguishing between a normal event and a cyberattack can be difficult because label information is mostly unavailable. Therefore, it is imperative to develop a threat intelligence system that operates more effectively without supervision, i.e., without a label. Additionally, reducing the false positive rate in cyber threat detection is a more promising step for a safer and more reliable environment. Also, the enormous number of features in the data for intrusion detection tasks sometimes results in significant computing costs. Therefore, a novel hybrid feature selection based unsupervised intrusion detection system is proposed, which is termed as HYbrid and Robust Intrusion DEtection (HYRIDE), that uses a wide variety of feature selection techniques to obtain the fewest, best possible features. The local outlier factor, elliptic envelope, and histogram-based outlier score models are then trained using these features to identify threats in network traffic automatically. As a result, HYRIDE can effectively and efficiently distinguish between normal events and intrusions. The proposed methodology is empirically evaluated using popular datasets such as Telemetry datasets of Internet of Things (IoT) services, Operating systems datasets of Windows and Linux, as well as datasets of Network traffic (TON_IoT), University of New South Wales-Network Benchmark (UNSW-NB15), and Canadian Institute of Cybersecurity Intrusion Detection System (CICIDS 2017).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".