This paper proposes a hybrid framework, VAE-Trans, for Chinese calligraphy generation. This method achieves decoupled representation of style and glyph structure through a conditional variational autoencoder and introduces a stroke-aware Transformer decoder to enhance attention modeling capabilities for key structural regions. Simultaneously, multi-objective joint training and KL annealing are employed to balance generation quality and distribution consistency. Experiments validate the effectiveness of the method on HCL2000, CASIA-HWDB, and a self-built few-shot calligraphy style dataset. Results show that this method outperforms existing methods in terms of content accuracy, style accuracy, and structure preservation, especially in complex Chinese characters with multiple radicals and few-shot style transfer scenarios.
Xu et al. (Sat,) studied this question.