LogMaster: Mining Event Correlations in Logs of Large-scale Cluster Systems

被引:33
|
作者
Fu, Xiaoyu [1 ]
Ren, Rui [1 ]
Zhan, Jianfeng [1 ]
Zhou, Wei [1 ]
Jia, Zhen [1 ]
Lu, Gang [1 ]
机构
[1] Chinese Acad Sci, Inst Comp Technol, State Key Lab Comp Architecture, Beijing, Peoples R China
来源
2012 31ST INTERNATIONAL SYMPOSIUM ON RELIABLE DISTRIBUTED SYSTEMS (SRDS 2012) | 2012年
关键词
reliability; correlation mining; failure prediction; large-scale systems;
D O I
10.1109/SRDS.2012.40
中图分类号
TP3 [计算技术、计算机技术];
学科分类号
0812 ;
摘要
This paper presents a set of innovative algorithms and a system, named LogMaster, for mining correlations of events that have multiple attributions, i.e., node ID, application ID, event type, and event severity, in logs of large-scale cloud and HPC systems. Different from traditional transactional data, e.g., supermarket purchases, system logs have their unique characteristics, and hence we propose several innovative approaches to mining their correlations. We parse logs into an n-ary sequence where each event is identified by an informative nine-tuple. We propose a set of enhanced apriori-like algorithms for improving sequence mining efficiency; we propose an innovative abstraction-event correlation graphs (ECGs) to represent event correlations, and present an ECGs-based algorithm for fast predicting events. The experimental results on three logs of production cloud and HPC systems, varying from 433490 entries to 4747963 entries, show that our method can predict failures with a high precision and an acceptable recall rates.
引用
收藏
页码:71 / 80
页数:10
相关论文
共 50 条
  • [1] Decentralized event-triggered control of large-scale nonlinear systems
    Liu, Tengfei
    Jiang, Zhong-Ping
    Zhang, Pengpeng
    INTERNATIONAL JOURNAL OF ROBUST AND NONLINEAR CONTROL, 2020, 30 (04) : 1451 - 1466
  • [2] Wireless ventilation control for large-scale systems: The mining industrial case
    Witrant, E.
    D'Innocenzo, A.
    Sandou, G.
    Santucci, F.
    Di Benedetto, M. D.
    Isaksson, A. J.
    Johansson, K. H.
    Niculescu, S-I
    Olaru, S.
    Serra, E.
    Tennina, S.
    Tiberi, U.
    INTERNATIONAL JOURNAL OF ROBUST AND NONLINEAR CONTROL, 2010, 20 (02) : 226 - 251
  • [3] Event-triggered distributed predictive control for constrained large-scale linear systems
    Su Xu
    Zou Yuanyuan
    Niu Yugang
    Jia Tinggang
    PROCEEDINGS OF THE 35TH CHINESE CONTROL CONFERENCE 2016, 2016, : 4253 - 4258
  • [4] Monitor Placement for Large-Scale Systems
    Talele, Nirupama
    Teutsch, Jason
    Erbacher, Robert
    Jaeger, Trent
    PROCEEDINGS OF THE 19TH ACM SYMPOSIUM ON ACCESS CONTROL MODELS AND TECHNOLOGIES (SACMAT'14), 2014, : 29 - 40
  • [5] Paradigm of Inheritance in Large-Scale Systems
    Gadasin, D., V
    Shvedov, A., V
    Litvin, Ya S.
    2019 SYSTEMS OF SIGNALS GENERATING AND PROCESSING IN THE FIELD OF ON BOARD COMMUNICATIONS, 2019,
  • [6] Event-Triggered Output-Feedback Control for Large-Scale Systems With Unknown Hysteresis
    Cao, Liang
    Ren, Hongru
    Li, Hongyi
    Lu, Renquan
    IEEE TRANSACTIONS ON CYBERNETICS, 2021, 51 (11) : 5236 - 5247
  • [7] Decentralized event-triggered output feedback control for a class of interconnected large-scale systems
    Liu, Dan
    Yang, Guang-Hong
    ISA TRANSACTIONS, 2019, 93 : 156 - 164
  • [8] Distributed Event-Triggered Control of Large-Scale Fuzzy Systems With Long Transmission Delays
    Cui, Hailong
    Zhao, Guanglei
    Hua, Changchun
    Cui, Ruixue
    IEEE TRANSACTIONS ON FUZZY SYSTEMS, 2024, 32 (06) : 3765 - 3778
  • [9] Sentiment Analysis based Error Detection for Large-Scale Systems
    Alharthi, Khalid Ayedh
    Jhumka, Arshad
    Di, Sheng
    Cappello, Franck
    Chuah, Edward
    51ST ANNUAL IEEE/IFIP INTERNATIONAL CONFERENCE ON DEPENDABLE SYSTEMS AND NETWORKS (DSN 2021), 2021, : 237 - 249
  • [10] Event-Triggered Robust Adaptive Dynamic Programming With Output Feedback for Large-Scale Systems
    Zhao, Fuyu
    Gao, Weinan
    Liu, Tengfei
    Jiang, Zhong-Ping
    IEEE TRANSACTIONS ON CONTROL OF NETWORK SYSTEMS, 2023, 10 (01): : 63 - 74