Neural Network Based Android Malware Detection with Different IP Coding Methods
Bibliographic record
Abstract
Due to the COVID-19 epidemic that has affected the \nwhole world, internet use has increased more than in previous \nyears. Almost all operations and transactions are done over the \ninternet, especially with the use of cellular phones and tablet PCs. \nThis growth results in many security deficits that need to be solved \nby security admins and end users. Malicious software (malware) \nis generally preferred for attacking the computer systems and \nrecently for cellular phones. As a mobile operating system, \nAndroid is the main player of this sector with about 72% market \nshare worldwide. Therefore, malware attacks especially target \nthese devices, for reaching the maximum number of victims. The \nsituation is getting more and more devastating with around 12,000 \nnew Android malware attacks every day. This is one critical \nproblem that needed to be solved by setting up an android \nmalware detection system. Machine learning algorithms are \nfrequently preferred in data mining-based security applications \nwhich contain lots of features in datasets. Artificial Neural \nnetworks are one of the mostly preferred learning models for \ntraining the system. Therefore, in this paper, it is aimed to \nimplement a neural network based android malware detection \nsystem by using an up-to-date dataset presented by the Cyber \nSecurity Institute of Canada as CICMalDroid2017. Ip Addresses \nare one of the features in this dataset, and we focus on two \ndifferent IP coding methods, as IP Splitting to Four Numbers, IP \nTransform to integer number, and no IP Address. In experimental \nstudy we reached a good level of accuracy rate as 98.4% by \nsplitting an IP address to four numbers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".