FADO: Floorplan-Aware Directive Optimization Based on Synthesis and Analytical Models for High-Level Synthesis Designs on Multi-Die FPGAs

被引：0

作者：

Du, Linfeng ^{[1
]}

Liang, Tingyuan ^{[1
]}

Zhou, Xiaofeng ^{[1
]}

Ge, Jinming ^{[1
]}

Li, Shangkun ^{[2
]}

Sinha, Sharad ^{[3
]}

Zhao, Jieru ^{[4
]}

Xie, Zhiyao ^{[1
]}

Zhang, Wei ^{[1
]}

机构：

[1] Hong Kong Univ Sci & Technol, Kowloon, Elect & Comp Engn, Hong Kong, Peoples R China

[2] Fudan Univ, Shanghai, Peoples R China

[3] Indian Inst Technol Goa, Comp Sci & Engn, Ponda, Goa, India

[4] Shanghai Jiao Tong Univ, Comp Sci & Engn, Shanghai, Peoples R China

来源：

ACM TRANSACTIONS ON RECONFIGURABLE TECHNOLOGY AND SYSTEMS | 2024年 / 17卷 / 03期

关键词：

High-level synthesis; analytical model; design space exploration; multi-die FPGA; directive optimization; floorplanning;

D O I：

10.1145/3653458

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

Multi-die FPGAs are widely adopted for large-scale accelerators, but optimizing high-level synthesis designs on these FPGAs faces two challenges. First, the delay caused by die-crossing nets creates an NP-hard floor- planning problem. Second, traditional directive optimization cannot consider resource constraints on each die or the timing issue incurred by the die-crossings. Furthermore, the high algorithmic complexity and the large scale lead to extended runtime for legalizing the floorplan of HLS designs under different directive configurations. To co-optimize the directives and floorplan of HLS designs on multi-die FPGAs, we formulate the co-search based on bin-packing variants and present two iterative optimization flows. The first (FADO 1.0) relies on a pre-built QoR library. It involves a greedy, latency-bottleneck-guided directive search, and an incremental floorplan legalization. Compared with a global floorplanning solution, it takes 693X similar to 4925X similar to 4925X shorter search time and achieves 1.16X similar to 8.78X similar to 8.78X better design performance, measured in workload execution time. To remove the time-consuming QoR library generation, the second flow (FADO 2.0) integrates an analytical QoR model and redesigns the directive search to accelerate convergence. Through experiments on mixed dataflow and non-dataflow designs, compared with 1.0, FADO 2.0 further yields a 1.40X better design performance on average after implementation on the Alveo U250 FPGA.

引用

页数：33

共 35 条

[31] Simulated annealing-based high-level synthesis methodology for reliable and energy-aware application specific integrated circuit designs with multiple supply voltages
Dilek, Selma
Tosun, Suleyman
Cakin, Alperen
INTERNATIONAL JOURNAL OF CIRCUIT THEORY AND APPLICATIONS, 2023, 51 (10) : 4897 - 4938
[32] Hotspot Mitigation through Multi-Row Thermal-aware Re-Placement of Logic Cells based on High-Level Synthesis Scheduling
Schafer, Benjamin Carrion
27TH ASIA AND SOUTH PACIFIC DESIGN AUTOMATION CONFERENCE, ASP-DAC 2022, 2022, : 629 - 634
[33] SoC-based solution for multi-axis control systems using high-level synthesis
Gu, Qiang
Li, Yesong
Niu, Panqing
PROCEEDINGS OF THE INSTITUTION OF MECHANICAL ENGINEERS PART I-JOURNAL OF SYSTEMS AND CONTROL ENGINEERING, 2015, 229 (01) : 63 - 73
[34] Flexible design of wide-pipeline-based WiMAX QC-LDPC decoder architectures on FPGAs using high-level synthesis
Andrade, J.
Falcao, G.
Silva, V.
ELECTRONICS LETTERS, 2014, 50 (11) : 839 - U92
[35] Bit-length optimization method for high-level synthesis based on non-linear programming technique
Doi, Nobuhiro
Horiyama, Takashi
Nakanishi, Masaki
Kimura, Shinji
IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES, 2006, E89A (12) : 3427 - 3434

← 1 2 3 4 →