PreGAN: Preemptive Migration Prediction Network for Proactive Fault-Tolerant Edge Computing

被引:33
作者
Tuli, Shreshth [1 ]
Casale, Giuliano [1 ]
Jennings, Nicholas R. [1 ,2 ]
机构
[1] Imperial Coll London, London, England
[2] Loughborough Univ, Loughborough, Leics, England
来源
IEEE CONFERENCE ON COMPUTER COMMUNICATIONS (IEEE INFOCOM 2022) | 2022年
关键词
Fault Tolerance; Preemptive Migrations; Edge Computing; Generative Adversarial Networks; STRATEGY;
D O I
10.1109/INFOCOM48880.2022.9796778
中图分类号
TP3 [计算技术、计算机技术];
学科分类号
0812 ;
摘要
Building a fault-tolerant edge system that can quickly react to node overloads or failures is challenging due to the unreliability of edge devices and the strict service deadlines of modern applications. Moreover, unnecessary task migrations can stress the system network, giving rise to the need for a smart and parsimonious failure recovery scheme. Prior approaches often fail to adapt to highly volatile workloads or accurately detect and diagnose faults for optimal remediation. There is thus a need for a robust and proactive fault-tolerance mechanism to meet service level objectives. In this work, we propose PreGAN, a composite AI model using a Generative Adversarial Network (GAN) to predict preemptive migration decisions for proactive fault-tolerance in containerized edge deployments. PreGAN uses co-simulations in tandem with a GAN to learn a few-shot anomaly classifier and proactively predict migration decisions for reliable computing. Extensive experiments on a Raspberry-Pi based edge environment show that PreGAN can outperform state-of-the-art baseline methods in fault-detection, diagnosis and classification, thus achieving high quality of service. PreGAN accomplishes this by 5.1% more accurate fault detection, higher diagnosis scores and 23.8% lower overheads compared to the best method among the considered baselines.
引用
收藏
页码:670 / 679
页数:10
相关论文
共 46 条
[31]   Fault-tolerant with load balancing scheduling in a fog-based IoT application [J].
Sharif, Ahmad ;
Nickray, Mohsen ;
Shahidinejad, Ali .
IET COMMUNICATIONS, 2020, 14 (16) :2646-2657
[32]   An Improved Dynamic Fault Tolerant Management Algorithm during VM migration in Cloud Data Center [J].
Sivagami, V. M. ;
Easwarakumar, K. S. .
FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2019, 98 :35-43
[33]  
Snell J, 2017, ADV NEUR IN, V30
[34]  
Standard Performance Evaluation Corporation, SPEC POW CONS MOD
[35]   Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network [J].
Su, Ya ;
Zhao, Youjian ;
Niu, Chenhao ;
Liu, Rong ;
Sun, Wei ;
Pei, Dan .
KDD'19: PROCEEDINGS OF THE 25TH ACM SIGKDD INTERNATIONAL CONFERENCCE ON KNOWLEDGE DISCOVERY AND DATA MINING, 2019, :2828-2837
[36]  
Tian BC, 2018, IEEE INFOCOM SER, P864, DOI 10.1109/INFOCOM.2018.8486340
[37]  
Tuli S., 2021, INT JOINT C ARTIFICI
[38]  
Tuli S., 2019, J SYST SOFTWARE
[39]  
Tuli S., 2021, J SYST SOFTWARE, P111
[40]   COSCO: Container Orchestration Using Co-Simulation and Gradient Based Optimization for Fog Computing Environments [J].
Tuli, Shreshth ;
Poojara, Shivananda R. ;
Srirama, Satish N. ;
Casale, Giuliano ;
Jennings, Nicholas R. .
IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, 2022, 33 (01) :101-116