Android botnet detection using signature \ndata and Ensemble Machine Learning
Bibliographic record
Abstract
As the use of smartphones has increased intensely in the past decade for daily activities such as socialising, banking, online shopping and communicating with friends and family. Android operating system is very popular and used universally for smartphones and tablets. Therefore, threats for this android platform is emerging very rapidly. Exploiting smartphones are comparatively easy and more effective than exploiting traditional computer systems and thus attackers started developing applications with hidden botnet capabilities. These applications use to take control of user’s device without his permission to steal sensitive data or launch denial-of-service attack with the help of Command and Control (C&C) servers. There are many proposed solutions available to detect botnet application using various approaches. In this paper, I proposed a hybrid model for botnet detection using a combination of signature-based detection at initial layer to perform abrupt detection. At 2nd layer ensemble machine learning method is used to identify botnet components with the help of extracted permissions and intents via static analysis. I compared 5 machine learning classifier algorithms and selected three with highest accuracy to create ensemble model. To extract the features to prepare efficient dataset for training and testing of this machine learning model I analyse 375 applications with botnet capabilities and 1105 benign applications from CICInvesAndMal2019 dataset which is novel and publicly available for researchers by the Canadian Institute for Cybersecurity. To confirm this result, we used Virus Total as a reference point which also showed comparable results of botnet detection. In this experiment, we successfully obtain 95.4% accuracy with the Logistic Regression classifier which was slightly increased to 95.8% after assembling top three algorithms. \nKeywords: Android Botnets, Ensemble Machine Learning, Signature-Based detection, Permissions, Intents, and DDoS prevention.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".