malware-detectionsupply-chain-securitynpm-ecosystemmachine-learning

Sejfia & Schäfer — Amalfi, ICSE

it directly studies the detection of malicious npm packages using features that overlap with Risk Guard's check signals -- install scripts, source repository availability, package-to-source reproducibility, publication timing patterns, and metadata anomalies -- providing empirical evidence for how these signals distinguish malicious from benign packages in the npm ecosystem.

Summary

AMALFI is a machine-learning system for automated detection of malicious npm packages, evaluated at ICSE 2022. It combines three classifiers (decision tree, Naive Bayes, one-class SVM) with a source-code reproducer and a clone detector. Trained on a corpus of 643 known malicious and 1,147 benign npm package versions provided by npm, it was run against 96,287 newly published package versions over one week (July 29 to August 4, 2021) and identified 95 previously unknown malware samples, all confirmed by npm's security team. The system extracts 11 features including install script presence, network access, file-system access, dynamic code generation (eval/Function), data encoding, cryptographic API usage, PII access patterns, minified/binary file entropy, time between updates, and semantic version update type. Of the 95 malicious packages found, 33 used install scripts, 52 were major version updates, 30 were first-version publishes, 11 exhibited file-system access, and 10 exhibited network access. Malicious packages showed median entropy of 4.69 versus 0.001 for benign, and median time between updates of 7.02 seconds versus 2,217.18 seconds for benign packages. The decision tree classifier achieved 0.98 precision and 0.43 recall on cross-validation; the reproducer had low success rates but eliminated false positives by verifying packages could be rebuilt from source; the clone detector identified verbatim copies of known malware via MD5 hashing of tarball contents (excluding package name/version). The system processes feature extraction in under 10 seconds per package and classification in under 1 second.

Related Checks

PACKAGE_INSTALL_SCRIPTS

33 of 95 confirmed malicious packages used install scripts (preinstall/install/postinstall) as their primary attack vector, including mogodb which used a postinstall script to exfiltrate hostnames to an attacker-controlled server

Adverse Outcome

arbitrary code execution on consumer machines during package installation, enabling credential theft, data exfiltration, and cryptomining before any application code runs

Because

AMALFI's feature analysis confirmed install scripts as one of the strongest single predictors of malicious intent, appearing in 35% of detected malware, and the decision tree classifiers used install script presence as a primary splitting feature across all seven days of evaluation

SOURCE_REPO_NOT_FOUND

AMALFI's reproducer component found that malicious packages tend not to have publicly available source code to avoid detection; packages that could be reproduced from source were filtered as benign, and the reproducer's low success rate (1-2 packages per day) reflected how few flagged packages had accessible source repositories

Adverse Outcome

inclusion of packages whose contents cannot be independently verified against source code, preventing detection of injected malicious payloads

Because

the paper establishes that source repository absence is a reliable negative signal -- malicious actors deliberately omit or provide inaccessible repository URLs to prevent source-to-package comparison that would reveal inserted malware

PACKAGE_RELEASE_COOLDOWN

Malicious packages exhibited a median time between updates of 7.02 seconds compared to 2,217.18 seconds for benign packages, and the mogodb typosquatting attack published two versions within less than a millisecond

Adverse Outcome

installing a recently published malicious package version before community vetting or security scanning can detect the compromise

Because

the paper's entropy and timing analysis demonstrated that abnormally short publication intervals are a statistically significant discriminator between malicious and benign packages, with malicious median update time over 300x shorter than benign

PACKAGE_PAST_MALWARE

The jasmin package was benign in versions 0.0.1 and 0.0.2 but compromised in version 0.0.3 via account takeover, demonstrating that packages with prior malware history carry elevated ongoing risk; AMALFI's corpus contained 63 compromised versions of otherwise legitimate packages

Adverse Outcome

continued dependency on a package with demonstrated susceptibility to maintainer compromise or malicious code injection

Because

the paper documents that compromised packages like eslint-scope and event-stream were installed millions of times before detection, and that packages compromised once face elevated risk because attackers specifically target packages with known security governance weaknesses

PACKAGE_ACTIVE_MALWARE

AMALFI identified 95 previously unknown actively malicious packages in a single week, all confirmed and taken down by npm's security team, demonstrating the scale of active malware in the npm registry

Adverse Outcome

installing and executing packages containing active malicious code that exfiltrates credentials, harvests PII, or performs unauthorized computations

Because

the paper's week-long scan of 96,287 package versions proved that active malware is continuously published to npm at a rate that makes manual review infeasible, validating automated detection checks as essential for supply chain protection

SOURCE_SINGLE_CONTRIBUTOR

The event-stream attack succeeded because a single maintainer was socially engineered into granting access, and the jasmin compromise exploited a sole maintainer's stolen credentials; AMALFI notes that single-maintainer packages lack independent code review

Adverse Outcome

supply chain compromise through social engineering or credential theft targeting the sole maintainer of a package

Because

the paper documents multiple real-world incidents where attackers specifically targeted single-maintainer packages because there was no second reviewer to detect malicious code insertion, making sole maintainership a reliable proxy for social engineering vulnerability

Gaps Analysis

Evidence

Malicious packages exhibited median update intervals of 7.02 seconds versus 2,217.18 seconds for benign packages, and multiple malicious versions (e.g., mogodb 3.1.8 and 3.1.9) were published within milliseconds of each other

Blind Spot

Risk Guard's PACKAGE_RELEASE_COOLDOWN uses a 7-day threshold but does not measure abnormally rapid successive version publishing (sub-minute intervals) as a distinct malware indicator

Actionable Capability

Risk Guard would be better if it flagged packages with multiple versions published within minutes or seconds of each other as a high-confidence malware signal, separate from the general cooldown check

Evidence

52 of 95 malicious packages were major semver updates and 30 were first-version publishes, showing that update type strongly correlates with malicious intent when combined with other signals

Blind Spot

Risk Guard does not analyze the semantic version update type (major/minor/patch/prerelease/first) as a risk signal or correlate it with other metadata anomalies

Actionable Capability

Risk Guard would be better if it tracked version update type and flagged unexpected capability introductions (e.g., install scripts appearing for the first time in a minor version bump)

Evidence

Malicious packages had median file entropy of 4.69 versus 0.001 for benign packages, indicating presence of minified code or binary executables used to evade detection

Blind Spot

Risk Guard does not measure file entropy or detect the presence of obfuscated, minified, or binary executable content in published packages

Actionable Capability

Risk Guard would be better if it flagged packages containing unusually high-entropy files (minified code, embedded binaries) that are inconsistent with the package's stated purpose

← Previous Next →