ABSTRACT Deep learning has shown strong potential for image‐based aesthetic assessment, but its limited interpretability restricts its use in design‐related applications. This paper proposes XAesViT‐Net, an explainable deep learning framework for car front‐end aesthetic assessment. Inspired by visual aesthetic cognition, the proposed network integrates local feature extraction, global feature modelling and attention‐based feature fusion to improve predictive performance while introducing an interpretability‐oriented architectural design. To evaluate the method, an auto face aesthetics dataset (AFAD) of car front‐view images was constructed using publicly available platform ratings and the task was formulated as a three‐class aesthetic classification. Experimental results show that XAesViT‐Net achieves superior performance against both generic backbones and representative task‐specific aesthetic assessment methods on AFAD, reaching 96.34% accuracy and 0.9972 AUC‐ROC. In addition, Grad‐CAM and SHAP were employed to identify the visual regions and design cues underlying the model predictions, providing post‐hoc interpretability analyses that improve transparency and practical usability. The results demonstrate that the proposed framework is effective for interpretable car front‐end aesthetic assessment. The current validation is limited to car front‐end images and extension to broader product categories remains future work.
Ouyang et al. (Thu,) studied this question.