PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 26, 20240 citations

Post-Training Quantization of CNN Accelerator with Mixed Precision Floating-Point Arithmetic Selection Based on a Genetic Algorithm

View Full Paper
MJMuhammad JunaidHAHayotjon AlievSSShoaib Sajid

Key Points

Key points are not available for this paper at this time.

Abstract

The increasing complexity of Deep Neural Network (DNN) models, coupled with rising demands for energy efficiency and computational speed in edge devices, necessitates the reevaluation of conventional numerical representations and computational strategies. Traditional quantization of DNNs, which often employs low-bit integers due to their fixed hardware implementation, applies uniform granularity across all values. This approach can be inefficient when certain data points require finer granularity to maintain accuracy. In contrast, floating-point (FP) quantization provides greater flexibility in bit allocation, enabling more precise granularity where needed. This paper proposes mixed-precision FP arithmetic based on a genetic algorithm to find the optimal precision for each layer. To minimize rounding errors, stochastic rounding is incorporated, and layer-specific exponent bias adjustments are implemented to enhance data representation precision. Experimental results from implementing the proposed mixed-precision FP quantization on the YOLOv2-tiny model reveal a 2.9 times reduction in energy consumption per image and a 50% reduction in memory requirements, with only a negligible loss of 0.13% in mean Average Precision (mAP@0.5) as evaluated on the VOC dataset, compared to Bfloat16, a known low-cost floating-point method.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Junaid et al. (2024) studied this question.

synapsesocial.com/papers/68e5ef77b6db643587583db6https://doi.org/10.36227/techrxiv.172202612.21470738/v1
Ask AI
Helpful
Bookmark
Share
View Full Paper