PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 23, 2026International Journal of Circuit Theory and Applications0 citations

Enhanced DSP Architecture for Small Floating‐Point Based Deep Learning Accelerators on FPGAs

View Full Paper
KMKuiming MaCLChaoying LiuYZYang Zhang

Key Points

  • This work aims to improve computational efficiency in deep learning models by enhancing DSP architecture on FPGAs.
  • Decomposed a 27 18-bit multiplier into four 8 18-bit multipliers.
  • Implemented left-shift operations using one-hot encoding for number conversion.
  • Compared synthesis results with standard DSP48E2 to evaluate performance.
  • Increased compute density for INT8 MACs by 129%, 106%, and 120% for different kernel sizes.
  • Achieved a 10% area cost for supporting 4 INT8-based MACs.
  • Reduced LUT6s in the process element array by two-thirds for 8-bit floating-point convolution cores.

Abstract

ABSTRACT Deep learning models should be quantized as the low bit‐width representation for improving computational efficiency. However, the digital signal processing (DSP) blocks in the field‐programmable gate array cannot achieve high efficiency for the low bit‐width fixed‐point based multiply‐accumulate operations (MACs). In addition, the low bit‐width floating‐point based deep learning models are widely used in recent years that results in a new challenge for DSP blocks. Therefore, this brief decomposes the large 27 18‐bit multiplier in the DSP48E2 into four 8 18‐bit multipliers to increase the compute density of INT8‐based MACs and uses these 8 18 multipliers to perform left‐shift operations with the one‐hot encoding scheme that facilitates the conversion from the floating‐point numbers to the fixed‐point numbers. In contrast with the DSP48E2, the synthesis results prove that the enhanced DSP48E2 supports 4 INT8‐based MACs at the cost of 10% area and converts 4 low bit‐width floating‐point numbers as the fixed‐point numbers at the cost of an additional 1% area. In addition, for the 8‐bit fixed‐point based convolution cores, the number of MACs for each enhanced DSP48E2 is improved by 129%, 106%, and 120% for , 3 3, and 5 5 kernels. For the 8‐bit floating‐point based convolution cores, the LUT6s of process element array can be reduced by two‐thirds.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ma et al. (2026) studied this question.

synapsesocial.com/papers/69c08bcaa48f6b84677f98ffhttps://doi.org/10.1002/cta.70387
Ask AI
Helpful
Bookmark
Share
View Full Paper