This study undertakes a case-based qualitative analysis of the performance of contemporary Artificial Intelligence (AI) models to translate lexical items from the Holy Qur’an that share a core denotation but possess distinct connotations based on their textual context. The Qur’an is distinguished by a sophisticated lexicon where such words, termed cognitive synonyms, carry discrete semantic loads contingent on their specific collocation. While machine translation has historically struggled to preserve these distinctions, this study evaluates the performance of three leading Large Language Models (LLMs): GPT-4, Gemini Pro 2.5, and DeepSeek-V3.2. Focusing on a purposively bounded corpus of sixteen lexical pairs spanning theological, anthropological, and cosmological domains, the paper qualitatively analyzes AI-generated translations. These outputs were examined in relation to Abdel Haleem’s English translation and selected insights from classical Islamic exegesis (tafsīr), treated as interpretive reference points rather than a unified exegetical framework. The analysis indicates that, within these sixteen examined cases and under the specified prompting conditions, current LLMs in several instances produced outputs that favored more general equivalents, which in some of the examples led to a reduction of context-specific semantic distinctions, rather than systematically preserving fine-grained meanings. Within this bounded dataset, the findings further suggest that while LLMs were able to recognize explicit contextual patterns, they did not consistently preserve finer semantic nuances. The study concludes that within the limited dataset examined, and without explicit hermeneutic guidance, general-purpose AI models in these tested cases may be less reliable for high-fidelity Qur’anic translation, functioning as powerful probabilistic tools rather than systems of full semantic understanding.
Shehab et al. (Sun,) studied this question.