ABSTRACT Deep learning models should be quantized as the low bit‐width representation for improving computational efficiency. However, the digital signal processing (DSP) blocks in the field‐programmable gate array cannot achieve high efficiency for the low bit‐width fixed‐point based multiply‐accumulate operations (MACs). In addition, the low bit‐width floating‐point based deep learning models are widely used in recent years that results in a new challenge for DSP blocks. Therefore, this brief decomposes the large 27 18‐bit multiplier in the DSP48E2 into four 8 18‐bit multipliers to increase the compute density of INT8‐based MACs and uses these 8 18 multipliers to perform left‐shift operations with the one‐hot encoding scheme that facilitates the conversion from the floating‐point numbers to the fixed‐point numbers. In contrast with the DSP48E2, the synthesis results prove that the enhanced DSP48E2 supports 4 INT8‐based MACs at the cost of 10% area and converts 4 low bit‐width floating‐point numbers as the fixed‐point numbers at the cost of an additional 1% area. In addition, for the 8‐bit fixed‐point based convolution cores, the number of MACs for each enhanced DSP48E2 is improved by 129%, 106%, and 120% for , 3 3, and 5 5 kernels. For the 8‐bit floating‐point based convolution cores, the LUT6s of process element array can be reduced by two‐thirds.
Ma et al. (2026) studied this question.