Optimum Interval for Application-level Checkpoints

被引:6
作者
Siavvas, Miltiadis [1 ,2 ]
Gelenbe, Erol [3 ]
机构
[1] Imperial Coll London, London, England
[2] Ctr Res & Technol Hellas, Thessaloniki, Greece
[3] Polish Acad Sci, Inst Theoret & Appl Informat, Gliwice, Poland
来源
2019 6TH IEEE INTERNATIONAL CONFERENCE ON CYBER SECURITY AND CLOUD COMPUTING (IEEE CSCLOUD 2019) / 2019 5TH IEEE INTERNATIONAL CONFERENCE ON EDGE COMPUTING AND SCALABLE CLOUD (IEEE EDGECOM 2019) | 2019年
基金
欧盟地平线“2020”;
关键词
Cloud Computing; Software Reliability; Roll Back Recovery; Application Level Checkpoints; Optimum Checkpoints; Program Loops; AVAILABILITY; SYSTEMS;
D O I
10.1109/CSCloud/EdgeCom.2019.000-4
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Checkpointing is commonly adopted for enhancing the performance of software applications that operate in the presence of failures. Among the existing checkpointing strategies, Application-level Checkpoint and Restart (ALCR) is considered the most efficient, since it leaves smaller memory footprint, but it requires significant development effort. Although existing ALCR tools and libraries manage to reduce the effort required for implementing the checkpoints, they do not provide recommendations regarding their inter-checkpoint interval. To this end, in the present paper, we develop a mathematical model to estimate the optimum checkpoint interval, i.e., the interval between two successive checkpoints that minimises the average execution time of the application. The case of programs with loops and nested loops is also discussed. The results are illustrated with several numerical examples.
引用
收藏
页码:145 / 150
页数:6
相关论文
共 40 条
[1]  
[Anonymous], TOP 20 HIGH PROFILE
[2]  
[Anonymous], 2018, SUMMARY AMAZON S3 SE
[3]  
[Anonymous], 1982, INTRO RESEAUX FILES
[4]  
[Anonymous], 1975, SCIENCE
[5]  
Ansel J., 2009, IPDPS 2009 P 2009 IE
[6]   A View of Cloud Computing [J].
Armbrust, Michael ;
Fox, Armando ;
Griffith, Rean ;
Joseph, Anthony D. ;
Katz, Randy ;
Konwinski, Andy ;
Lee, Gunho ;
Patterson, David ;
Rabkin, Ariel ;
Stoica, Ion ;
Zaharia, Matei .
COMMUNICATIONS OF THE ACM, 2010, 53 (04) :50-58
[7]  
Arora R., 2017, P 4 INT WORKSH HPC U
[8]   Checkpointing as a Service in Heterogeneous Cloud Environments [J].
Cao, Jiajun ;
Simonin, Matthieu ;
Cooperman, Gene ;
Morin, Christine .
2015 15TH IEEE/ACM INTERNATIONAL SYMPOSIUM ON CLUSTER, CLOUD AND GRID COMPUTING, 2015, :61-70
[9]   Automatic Workarounds: Exploiting the Intrinsic Redundancy of Web Applications [J].
Carzaniga, Antonio ;
Gorla, Alessandra ;
Perino, Nicolo ;
Pezze, Mauro .
ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY, 2015, 24 (03)
[10]  
Chen L., 1978, FTCS-8. The Eighth Annual International Conference on Fault-Tolerant Computing, P3