Ezieddin Elmahjub, Junaid Qadir, A. Mushtaq, Rafay Naeem, Ibrahim Ghaznavi, Waleed Iqbal · Artificial Intelligence and Law 2026 · 2026
DOI: 10.1007/s10506-026-09535-4
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As millions of Muslims worldwide turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question emerges: Can these AI systems reliably reason about Islamic law? This paper introduces , the first multi-school benchmark for evaluating LLM performance across a broad range of Islamic jurisprudence ( fiqh ). Drawing from 37 foundational legal texts spanning 1,200 years and seven schools of jurisprudence, we create 718 evaluation instances across 13 tasks, manually collected and organized by complexity, from basic recall to sophisticated reasoning including legal rationale identification ( ‘illah ), analogical application ( qiyās ), and cross-school synthesis. Our evaluation of nine state-of-the-art LLMs reveals significant limitations. Even the best model achieves only 67.65% correctness with 21.25% hallucination; several models achieved correctness below 35% and hallucination exceeding 55%. Furthermore, few-shot prompting yields no consistent benefit; per-model changes range from -3.64 to +4.16 points, with gains concentrated in already-weak models, consistent with insufficient Islamic legal knowledge in current training data that prompting, under closed-book conditions, could not compensate for . Our analysis reveals why: moderate-complexity tasks (from a human expert’s perspective) requiring exact, verbatim knowledge (e.g., enumerating contract conditions, synthesizing statutory articles) show consistently high error rates and hallucination rates up to 73%, whereas high-complexity tasks show better performance because models rely on semantic generalization and verbose reasoning, projecting competence while lacking precise textual understanding. False premise detection reveals risky sycophantic behavior: under few-shot prompting, 5 of 9 models accept misleading assumptions at rates exceeding 40% (worst: 86.27%). Critically, few-shot prompting worsens sycophancy by 3.49 percentage points. The strong and statistically significant negative correlation (Pearson’s r = –0.91, n = 9, p < 0.01) between the false premise acceptance rate and overall performance indicates that models with higher false Islamic query acceptance rates tend to exhibit lower overall reasoning accuracy. Our findings carry important implications: the path forward for Islamic NLP lies not in lightweight post-hoc techniques such as prompt or instruction tuning but in enriching knowledge. The Islamic NLP community must prioritize training models on comprehensive, large-scale Islamic legal corpora that span classical Hadith collections, jurisprudential works from the major Islamic law schools ( madhabs ), and codified legal compendia. provides the first systematic evaluation framework for Islamic legal AI systems, revealing critical limitations in platforms that Muslims increasingly rely on for spiritual guidance.
No comments yet — start the discussion below.