Leveraging Feature Selection and Deep Learning for Accurate Malware and Ransomware Detection in PE Files
Bibliographic record
Abstract
AI -driven malware detection, particularly using machine learning (ML) and deep learning (DL), has shown promise in the field of mal ware detection. While many studies focus on portable executable (PE) file features, fewer explore ransom ware detection using PE headers with deep learning. Additionally, limited research examines how feature selection impacts machine learning performance, which is crucial for optimizing detection accuracy. This paper investigates feature selection strategies to improve malware detection while minimizing feature count. We analyzed different dataset segments' impact on ML algorithms, refining a strategy to determine the optimal dataset proportion for training. We applied Principal Component Analysis, Mutual Information, and Chi-square feature selection techniques on two datasets: (1) 2,157 ransomware PE-header samples with 1,028 features and (2) 29,807 Windows malware samples with 54 features. Seven ML models were tested alongside deep learning models. For ransomware detection, the LSTM model achieved an accuracy of 99.7%. In the Malware dataset, an accuracy of 99.9% was obtained across all evaluation metrics using only 10 features selected with the Chi-square method when applied with NB, LR, ET, and SVM. Comparable results were achieved using mutual information in conjunction with RF and LR. For deep learning, the LSTM model attained an accuracy of 98.9%. In the Ransomware dataset, an accuracy of 99.9% was achieved using RF and ET with Chi-square on 500 selected features out of 1,027. PCA combined with LR resulted in an accuracy of 99.4%. These results emphasize the effectiveness of feature selection in enhancing both the accuracy and efficiency of mal ware and ransom ware detection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".