GeNIS: A modular dataset for network intrusion detection and classification

被引：0

作者：

Silva, Miguel ^{[1
]}

Pinto, Daniela ^{[1
]}

Vitorino, Joao ^{[1
]}

Goncalves, Jose ^{[1
]}

Maia, Eva ^{[1
]}

Praca, Isabel ^{[1
]}

机构：

[1] Polytech Porto ISEP IPP, Sch Engn, Res Grp Intelligent Engn & Comp Adv Innovat & Dev, P-4249015 Porto, Portugal

来源：

DATA IN BRIEF | 2025年 / 60卷

关键词：

Network flow; Packet capture; Attack classification; Anomaly detection; Machine learning; Cybersecurity; Dataset;

D O I：

10.1016/j.dib.2025.111487

中图分类号：

O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];

学科分类号：

07 ; 0710 ; 09 ;

摘要：

The development of artificial intelligence solutions for cyberattack detection and classification require high-quality and representative data. However, there is a scarcity of labelled datasets focused on the cyberattacks that target vulnerable small and medium-sized enterprises. To allow organizations to improve their intrusion detection systems according to their types of users, their active services, and the network protocols they use, it is necessary to provide reliable captures of different types of benign and malicious traffic. The GECAD Network Intrusion Scenarios (GeNIS) dataset contains multiple sequential attack scenarios and different types of realistic normal network activity, recorded during advanced network simulations on the Airbus CyberRange platform. The raw network packets were analyzed to generate labelled network flows, with the computation of statistical features to represent the traffic patterns of local and remote attackers, normal users and administrators, and background traffic of an enterprise computer network. GeNIS follows a modular design, providing raw packet capture next generation (PCAPNG) files with over 37 million packets of each intermediate attack step to enable an in-depth analysis with different flow exporters, feature extraction, and feature selection tools, as well as filtered CSV files with over 2.8 million flows created with 5, 10, 30, and 60 s flow intervals. The flows were preprocessed to provide a reliable benchmark dataset with the most relevant features for the training, validation, and testing of robust machine learning and deep learning models.

引用

页数：14