Simultaneous Botnet Attack Detection Using Long Short Term Memory-Based Autoencoder and XGBoost Classifier
Bibliographic record
Abstract
Botnet is a cyber-attack that aims to compromise the security of internet of things (IoT) networks by exploiting infected devices to launch attacks and gain unauthorized access to private information.Intrusion detection system (IDS) emerges as critical countermeasures for tackling the risks posed by botnet attacks, playing a crucial role in ensuring the integrity and confidentiality of data in IoT environments.Developing an effective botnet detection system depends on efficient contextual understanding and accurate attack pattern characterization.Recently, deep learning and machine learning based IDS have demonstrated promising results in traffic pattern recognition and identification as normal or malicious from raw data.However, these approaches fail to detect simultaneous botnet attacks as it ignores its distributed nature.In this paper, we propose an efficient hybrid deep learning model for simultaneous botnet attack detection over IoT networks.The twostage hybrid model analyzes the network traffic data captured from three parallel sensors and extracts simultaneous characteristics of attack traffic.The use of parallel detection enables more comprehensive coverage of the network, thereby increasing the detection accuracy of malicious activities that could be missed by a single sensor.Features are extracted using a long-short-term memory base autoencoder (LSTM-AE) over the NCC-2 Simultaneous Botnet Dataset.The LSTM-AE is trained using data from multiple sensors to model temporal characteristics and results in reduced latent representation.Attack type identification is achieved through a multi-class classification using the Extreme Gradient Boosting (XGBoost) ensemble learning algorithm.The recently released NCC-2 dataset is the first dataset to provide data representing sequential and simultaneous botnet activities detected concurrently by multiple sensors.Performance exploration indicates that for parallel botnet detection, the proposed LSTM-AE-XGB model achieves high accuracy while reducing false or missing detection.Moreover, to demonstrate model efficiency, we conducted a 10-fold cross-validation and a comparative performance analysis with the state-of-the-art ML and DL-based techniques for feature extraction and simultaneous botnet detection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".