Empathy—the ability to understand and respond to others’ emotions and perspectives—is a key communication skill for humans; however, it is under-explored within current conversational systems. While large language models (LLMs) have demonstrated a remarkable capability to generate coherent and contextually relevant output, they often struggle to exhibit genuine empathy, resulting in artificial and dull responses, particularly in low-resource languages such as Arabic. Notably, the research on empathetic conversational systems in Arabic is still in its early stages, mainly due to the scarcity of open-domain conversational data. To address this gap, we introduce Arabic Empathetic Conversations (AEConvs), a genuine Arabic conversational dataset featuring more than 4K open-domain dyadic empathetic conversations. This dataset provides a valuable resource that captures nuanced emotional and empathetic cues in the Arabic language. Using AEConvs, we evaluate and compare the empathetic capabilities of two state-of-the-art generative Arabic LLMs—AceGPT-chat and Jais-chat—under zero-shot and fine-tuning training settings. Human evaluation results demonstrate that while both models exhibit some form of empathy in zero-shot settings, fine-tuning on AEConvs improved their ability to generate more fine-grained empathetic responses while also yielding enhancements in fluency and context adherence. Additionally, automatic evaluation indicated improved language modeling and better lexical and semantic similarity with human reference responses. This study highlights the importance of culturally and linguistically tailored datasets in advancing empathetic conversational AI. We publicly release the AEConvs dataset, providing a valuable resource for future advancements in the field.
Alkhathlan et al. (Tue,) studied this question.