Evaluating LLM-Based Detection of Malicious Package Updates in npm
Bibliographic record
Abstract
The npm software package ecosystem is a notable target for adversarial actors, who seek to compromise software dependencies to exploit software developers and the end-users of their software. One especially dangerous form of attack involves the compromise of a package update. By sneaking malicious code into a package update, adversaries can trick package users into unknowingly installing malware. Detecting malicious package updates is an active research problem, as prospective solutions need to keep pace with the near-constant stream of new package updates, while also maintaining high detection accuracy. In this context, one potentially interesting and emergent approach involves utilizing large language models (LLMs) to identify malicious behaviors from the text of package code. However, practical use of LLMs also poses unique first-order challenges, as models are expensive to run and are known to struggle with task performance as input size increases. This work provides a critical exploration into the practicality and effectiveness of LLMs for detecting malicious package updates. We overcome the immediate challenges for LLM-based applications by preprocessing inputs for analysis and post-processing outputs for malware classification. We find this approach to be practical at repository scale and effective at detecting historical malware incidents, with our best-performing model correctly flagging 209 out of 209 malicious samples across a collection of historical attacks, while only flagging 8 out of 2,000 benign samples across a dataset of typical package updates. With first-order obstacles overcome, we then conduct a deeper investigation into the reasoning capabilities of LLMs–demonstrating specific mild code obfuscations that uniquely challenge tested LLMs and enable adaptive adversaries to subvert detection. Ultimately, our findings demonstrate nuanced potential for employing LLMs as a part of a larger security tool-belt for detecting package malware.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".