EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models

Anderson, Hyrum S.; Roth, Phil

Computer Science > Cryptography and Security

arXiv:1804.04637 (cs)

[Submitted on 12 Apr 2018 (v1), last revised 16 Apr 2018 (this version, v2)]

Title:EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models

Authors:Hyrum S. Anderson, Phil Roth

View PDF

Abstract:This paper describes EMBER: a labeled benchmark dataset for training machine learning models to statically detect malicious Windows portable executable files. The dataset includes features extracted from 1.1M binary files: 900K training samples (300K malicious, 300K benign, 300K unlabeled) and 200K test samples (100K malicious, 100K benign). To accompany the dataset, we also release open source code for extracting features from additional binaries so that additional sample features can be appended to the dataset. This dataset fills a void in the information security machine learning community: a benign/malicious dataset that is large, open and general enough to cover several interesting use cases. We enumerate several use cases that we considered when structuring the dataset. Additionally, we demonstrate one use case wherein we compare a baseline gradient boosted decision tree model trained using LightGBM with default settings to MalConv, a recently published end-to-end (featureless) deep learning model for malware detection. Results show that even without hyper-parameter optimization, the baseline EMBER model outperforms MalConv. The authors hope that the dataset, code and baseline model provided by EMBER will help invigorate machine learning research for malware detection, in much the same way that benchmark datasets have advanced computer vision research.

Subjects:	Cryptography and Security (cs.CR)
Cite as:	arXiv:1804.04637 [cs.CR]
	(or arXiv:1804.04637v2 [cs.CR] for this version)
	https://doi.org/10.48550/arXiv.1804.04637

Submission history

From: Hyrum Anderson [view email]
[v1] Thu, 12 Apr 2018 17:23:56 UTC (1,127 KB)
[v2] Mon, 16 Apr 2018 20:43:33 UTC (1,127 KB)

Computer Science > Cryptography and Security

Title:EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Cryptography and Security

Title:EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators