Based on an improved Loeffler architecture, a high-performance hardware accelerator for the 2D 8 x 8 Discrete Cosine Transform (DCT) and Inverse Discrete Cosine Transform (IDCT) is shown in this study. Designed for image and video processing applications, the suggested accelerator improves the Loeffler 8-point 1D DCT/IDCT data flow. By reducing the number of clock cycles needed for each operation and streamlining the arithmetic operations inside each cycle, an extremely effective 8-stage pipeline structure is used to increase processing performance. Fixed-point and canonic signed digit (CSD) coding methods are used to estimate the DCT coefficients using adders and shifters without the need for multiplication. Significantly lowering circuit complexity, the design includes a fast parallel transposed matrix architecture that effectively converts row-column coefficients. This architecture's FPGA implementation is demonstrated on the ARTIX-7 XC7VX330T chip, demonstrating its efficiency in high-speed image and video processing applications.
MOHANAIAH et al. (Thu,) studied this question.