fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms

被引：11

作者：

Chen, Yuedan ^{[1
,2
]}

Xiao, Guoqing ^{[1
,2
]}

Li, Kenli ^{[1
,2
]}

Piccialli, Francesco ^{[3
]}

Zomaya, Albert Y. ^{[4
]}

机构：

[1] Hunan Univ, Coll Comp Sci & Elect Engn, Changsha 410082, Hunan, Peoples R China

[2] Natl Supercomp Ctr Changsha, Changsha 410082, Hunan, Peoples R China

[3] Univ Naples Federico II, Dept Elect Engn & Informat Technol, I-80100 Naples, Italy

[4] Univ Sydney, Sch Informat Technol, Sidney, BC 2006, Canada

来源：

ACM TRANSACTIONS ON PARALLEL COMPUTING | 2022年 / 9卷 / 02期

基金：

中国国家自然科学基金;

关键词：

Heterogeneous; HPC; manycore; optimization; parallelism; SpMSpV; MATRIX-MATRIX MULTIPLICATION; PERFORMANCE;

D O I：

10.1145/3512770

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

Sparse matrix-sparse vector (SpMSpV) multiplication is one of the fundamental and important operations in many high-performance scientific and engineering applications. The inherent irregularity and poor data locality lead to two main challenges to scaling SpMSpV over high-performance computing (HPC) systems: (i) a large amount of redundant data limits the utilization of bandwidth and parallel resources; (ii) the irregular access pattern limits the exploitation of computing resources. This paper proposes a fine-grained parallel SpMSpV (fgSpMSpV) framework on Sunway TaihuLight supercomputer to alleviate the challenges for large-scale real-world applications. First, fgSpMSpV adopts an MPI+OpenMP+X parallelization model to exploit the multi-stage and hybrid parallelism of heterogeneous HPC architectures and accelerate both pre-/post-processing and main SpMSpV computation. Second, fgSpMSpV utilizes an adaptive parallel execution to reduce the pre-processing, adapt to the parallelism and memory hierarchy of the Sunway system, while still tame redundant and random memory accesses in SpMSpV, including a set of techniques like the fine-grained partitioner, re-collection method, and Compressed Sparse Column Vector (CSCV) matrix format. Third, fgSpMSpV uses several optimization techniques to further utilize the computing resources. fgSpMSpV on the Sunway TaihuLight gains a noticeable performance improvement from the key optimization techniques with various sparsity of the input. Additionally, fgSpMSpV is implemented on an NVIDIA Tesal P100 GPU and applied to the breath-first-search (BFS) application. fgSpMSpV on a P100 GPU obtains the speedup of up to 134.38x over the state-of-the-art SpMSpV algorithms, and the BFS application using fgSpMSpV achieves the speedup of up to 21.68x over the state-of-the-arts.

引用

页数：29

共 44 条

[21] PEPS plus plus : Towards Extreme-Scale Simulations of Strongly Correlated Quantum Many-Particle Models on Sunway TaihuLight [J].