Generating Arabic jurisprudential rulings requires precise reasoning and linguistic clarity, especially in sensitive domains such as Islamic inheritance. This study evaluates four experiments in large language models (LLMs) (RAG (Simple), RAG with Validation LLM, RAG with Re-ranking LLM, and Fine-Tuned LLM) within a three-layer evaluation framework combining equational, prompt-based, and human assessments. Results show that retrieval-augmented generation (RAG) approaches substantially outperform fine-tuned models across all metrics. The RAG (Validation LLM) achieved the highest overall performance with Recall = 0.78, F1 = 0.73, and BLEU = 0.30, surpassing the fine-tuned by more than 10–15%. In prompt-based evaluation, it scored above 0.8 in faithfulness and 0.9 in clearness, consistent with expert assessments that rated its clarity at 0.98 for definitions and 0.81 for inheritance cases. By basing the answers on original and reliable jurisprudential texts, this framework enhances factual reliability, interpretive transparency, and public accessibility, advancing explainable artificial intelligence for generating Arabic jurisprudential responses.
Mastor et al. (2026) studied this question.