Optimizing Multi-grid Computation and Parallelization on Multi-cores

被引:2
作者
Yang, Xiaojian [1 ]
Li, Shengguo [1 ]
Yuan, Fan [2 ]
Dong, Dezun [1 ]
Huang, Chun [1 ]
Wang, Zheng [3 ]
机构
[1] Natl Univ Def Technol, Changsha, Peoples R China
[2] Xiangtan Univ, Xiangtan, Peoples R China
[3] Univ Leeds, Leeds, W Yorkshire, England
来源
PROCEEDINGS OF THE 37TH INTERNATIONAL CONFERENCE ON SUPERCOMPUTING, ACM ICS 2023 | 2023年
基金
美国国家科学基金会; 国家重点研发计划;
关键词
Multigrid; symmetric Gauss-Seidel; Asynchronous parallelization; PERFORMANCE; EFFICIENT;
D O I
10.1145/3577193.3593726
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Multigrid algorithms are widely used to solve large-scale sparse linear systems, which is essential for many high-performance workloads. The symmetric Gauss-Seidel (SYMGS) method is often responsible for the performance bottleneck of MG. This paper presents new methods to parallelize and enhance the computation and parallelization efficiency of the SYMGS and MG algorithms on multi-core CPUs. Our solution employs a matrix splitting strategy and a revised computation formula to decrease the computation operations and memory accesses in SYMGS. With this new SYMGS strategy, we can then merge the two most time-consuming components of MG. On top of these, we propose a new asynchronous parallelization scheme to reduce the synchronization overhead when parallelizing SYMGS. We demonstrate the benefit of our techniques by integrating them with the HPCG benchmark and two real-life applications. Evaluation conducted on four architectures, including three ARMv8 and one x86, shows that our techniques greatly surpass the performance of engineer- and vendor-tuned implementations across various workloads and platforms.
引用
收藏
页码:227 / 239
页数:13
相关论文
共 62 条
  • [1] Anderson E., 1989, International Journal of High Speed Computing, V1, P73, DOI 10.1142/S0129053389000056
  • [2] [Anonymous], 2018, The HPCG benchmark: Analysis, shared memory preliminary improvements and evaluation on an arm-based platform
  • [3] Iterative Sparse Triangular Solves for Preconditioning
    Anzt, Hartwig
    Chow, Edmond
    Dongarra, Jack
    [J]. EURO-PAR 2015: PARALLEL PROCESSING, 2015, 9233 : 650 - 661
  • [4] Performance Optimization of the HPCG Benchmark on the Sunway TaihuLight Supercomputer
    Ao, Yulong
    Yang, Chao
    Liu, Fangfang
    Yin, Wanwang
    Jiang, Lijuan
    Sun, Qiao
    [J]. ACM TRANSACTIONS ON ARCHITECTURE AND CODE OPTIMIZATION, 2018, 15 (01)
  • [5] ARM, 2022, ARM performance libraries
  • [6] ARM Software, 2019, HPCG for ARM
  • [7] Assuncao J., 2017, P 15 INT C BRAZ GEOP, P1630
  • [8] Balay S., 2022, ANL-21/39-Revision3.17
  • [9] Communication lower bounds and optimal algorithms for numerical linear algebra
    Ballard, G.
    Carson, E.
    Demmel, J.
    Hoemmen, M.
    Knight, N.
    Schwartz, O.
    [J]. ACTA NUMERICA, 2014, 23 : 1 - 155
  • [10] MINIMIZING COMMUNICATION IN NUMERICAL LINEAR ALGEBRA
    Ballard, Grey
    Demmel, James
    Holtz, Olga
    Schwartz, Oded
    [J]. SIAM JOURNAL ON MATRIX ANALYSIS AND APPLICATIONS, 2011, 32 (03) : 866 - 901