Anomaly Detection and Anticipation in High Performance Computing Systems

被引:22
作者
Borghesi, Andrea [1 ,2 ]
Molan, Martin [1 ]
Milano, Michela [1 ,2 ]
Bartolini, Andrea [1 ,2 ]
机构
[1] Univ Bologna, DISI & DEI Dept, I-40126 Bologna, Italy
[2] Alma Mater Res Ctr Human Ctr Artificial Intellige, I-40126 Bologna, Italy
基金
欧盟地平线“2020”;
关键词
Supercomputers; Data models; Monitoring; Anomaly detection; Tools; Bridges; Computational modeling; High performance computing; anomaly detection; deep learning; DIAGNOSIS; NETWORK;
D O I
10.1109/TPDS.2021.3082802
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
In their quest toward Exascale, High Performance Computing (HPC) systems are rapidly becoming larger and more complex, together with the issues concerning their maintenance. Luckily, many current HPC systems are endowed with data monitoring infrastructures that characterize the system state, and whose data can be used to train Deep Learning (DL) anomaly detection models, a very popular research area. However, the lack of labels describing the state of the system is a wide-spread issue, as annotating data is a costly task, generally falling on human system administrators and thus does not scale toward exascale. In this article we investigate the possibility to extract labels from a service monitoring tool (Nagios) currently used by HPC system administrators to flag the nodes which undergo maintenance operations. This allows to automatically annotate data collected by a fine-grained monitoring infrastructure; this labelled data is then used to train and validate a DL model for anomaly detection. We conduct the experimental evaluation on a tier-0 production supercomputer hosted at CINECA, Bologna, Italy. The results reveal that the DL model can accurately detect the real failures, and, moreover, it can predict the insurgency of anomalies, by systematically anticipating the actual labels (i.e., the moment when system administrators realize when an anomalous event happened); the average advance time computed on historical traces is around 45 minutes. The proposed technology can be easily scaled toward exascale systems to easy their maintenance.
引用
收藏
页码:739 / 750
页数:12
相关论文
共 39 条
[1]  
[Anonymous], 2021, IEEE Trans. Broadcast.
[2]  
[Anonymous], 2020, KAIROSDB FAST SCALAB
[3]  
Apache, 2019, AP CASS
[4]  
Barth Wolfgang., 2008, NAGIOS SYSTEM NETWOR, V2nd
[5]   Paving theWay Toward Energy-Aware and Automated Datacentre [J].
Bartolini, Andrea ;
Beneventi, Francesco ;
Borghesi, Andrea ;
Cesarini, Daniele ;
Libri, Antonio ;
Benini, Luca ;
Cavazzoni, Carlo .
PROCEEDINGS OF THE 48TH INTERNATIONAL CONFERENCE ON PARALLEL PROCESSING WORKSHOPS (ICPP 2019), 2019,
[6]  
Beneventi F, 2017, DES AUT TEST EUROPE, P1038, DOI 10.23919/DATE.2017.7927143
[7]   Cost-Aware Prediction of Uncorrected DRAM Errors in the Field [J].
Boixaderas, Isaac ;
Zivanovic, Darko ;
More, Sergi ;
Bartolome, Javier ;
Vicente, David ;
Casas, Marc ;
Carpenter, Paul M. ;
Radojkovic, Petar ;
Ayguade, Eduard .
PROCEEDINGS OF SC20: THE INTERNATIONAL CONFERENCE FOR HIGH PERFORMANCE COMPUTING, NETWORKING, STORAGE AND ANALYSIS (SC20), 2020,
[8]   A semisupervised autoencoder-based approach for anomaly detection in high performance computing systems [J].
Borghesi, Andrea ;
Bartolini, Andrea ;
Lombardi, Michele ;
Milano, Michela ;
Benini, Luca .
ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2019, 85 :634-644
[9]  
Borghesi A, 2019, AAAI CONF ARTIF INTE, P9428
[10]  
Borghesi A, 2019, 2019 IEEE INTERNATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE CIRCUITS AND SYSTEMS (AICAS 2019), P229, DOI [10.1109/AICAS.2019.8771527, 10.1109/aicas.2019.8771527]