IslamicLegalBench: Evaluating LLMs knowledge and reasoning of islamic law across 1,200 years of Islamic pluralist legal traditions

Abstract As millions of Muslims worldwide turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question emerges: Can these AI systems reliably reason about Islamic law? This paper introduces , the first multi-school benchmark for evaluating LLM performance across a broad range of Islamic jurisprudence ( fiqh ). Drawing from 37 foundational legal texts spanning 1,200 years and seven schools of jurisprudence, we create 718 evaluation instances across 13 tasks, manually collected and organized by complexity, from basic recall to sophisticated reasoning including legal rationale identification ( ‘illah ), analogical application ( qiyās ), and cross-school synthesis. Our evaluation of nine state-of-the-art LLMs reveals significant limitations. Even the best model achieves only 67.65% correctness with 21.25% hallucination; several models achieved correctness below 35% and hallucination exceeding 55%. Furthermore, few-shot prompting yields no consistent benefit; per-model changes range from -3.64 to +4.16 points, with gains concentrated in already-weak models, consistent with insufficient Islamic legal knowledge in current training data that prompting, under closed-book conditions, could not compensate for . Our analysis reveals why: moderate-complexity tasks (from a human expert’s perspective) requiring exact, verbatim knowledge (e.g., enumerating contract conditions, synthesizing statutory articles) show consistently high error rates and hallucination rates up to 73%, whereas high-complexity tasks show better performance because models rely on semantic generalization and verbose reasoning, projecting competence while lacking precise textual understanding. False premise detection reveals risky sycophantic behavior: under few-shot prompting, 5 of 9 models accept misleading assumptions at rates exceeding 40% (worst: 86.27%). Critically, few-shot prompting worsens sycophancy by 3.49 percentage points. The strong and statistically significant negative correlation (Pearson’s r = –0.91, n = 9, p < 0.01) between the false premise acceptance rate and overall performance indicates that models with higher false Islamic query acceptance rates tend to exhibit lower overall reasoning accuracy. Our findings carry important implications: the path forward for Islamic NLP lies not in lightweight post-hoc techniques such as prompt or instruction tuning but in enriching knowledge. The Islamic NLP community must prioritize training models on comprehensive, large-scale Islamic legal corpora that span classical Hadith collections, jurisprudential works from the major Islamic law schools ( madhabs ), and codified legal compendia. provides the first systematic evaluation framework for Islamic legal AI systems, revealing critical limitations in platforms that Muslims increasingly rely on for spiritual guidance.

Authors

Institutions

Publication Details

Journal
Artificial Intelligence and Law
Published
2026-09-17
DOI
https://doi.org/10.1007/s10506-026-09535-4
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

IslamicLegalBench: Evaluating LLMs knowledge and reasoning of islamic law across 1,200 years of Islamic pluralist legal traditions

Ezieddin Elmahjub, A. Mushtaq, Ibrahim Ghaznavi, Rafay Naeem et al.
Artificial Intelligence and Law
Topic Modeling
article

IslamicLegalBench: Evaluating LLMs knowledge and reasoning of islamic law across 1,200 years of Islamic pluralist legal traditions

Ezieddin Elmahjub, A. Mushtaq, Ibrahim Ghaznavi, Rafay Naeem, Junaid Qadir, Waleed Iqbal
article en

Abstract

Abstract As millions of Muslims worldwide turn to LLMs like GPT, Claude, and DeepSeek for religious guidance, a critical question emerges: Can these AI systems reliably reason about Islamic law? This paper introduces , the first multi-school benchmark for evaluating LLM performance across a broad range of Islamic jurisprudence ( fiqh ). Drawing from 37 foundational legal texts spanning 1,200 years and seven schools of jurisprudence, we create 718 evaluation instances across 13 tasks, manually collected and organized by complexity, from basic recall to sophisticated reasoning including legal rationale identification ( ‘illah ), analogical application ( qiyās ), and cross-school synthesis. Our evaluation of nine state-of-the-art LLMs reveals significant limitations. Even the best model achieves only 67.65% correctness with 21.25% hallucination; several models achieved correctness below 35% and hallucination exceeding 55%. Furthermore, few-shot prompting yields no consistent benefit; per-model changes range from -3.64 to +4.16 points, with gains concentrated in already-weak models, consistent with insufficient Islamic legal knowledge in current training data that prompting, under closed-book conditions, could not compensate for . Our analysis reveals why: moderate-complexity tasks (from a human expert’s perspective) requiring exact, verbatim knowledge (e.g., enumerating contract conditions, synthesizing statutory articles) show consistently high error rates and hallucination rates up to 73%, whereas high-complexity tasks show better performance because models rely on semantic generalization and verbose reasoning, projecting competence while lacking precise textual understanding. False premise detection reveals risky sycophantic behavior: under few-shot prompting, 5 of 9 models accept misleading assumptions at rates exceeding 40% (worst: 86.27%). Critically, few-shot prompting worsens sycophancy by 3.49 percentage points. The strong and statistically significant negative correlation (Pearson’s r = –0.91, n = 9, p < 0.01) between the false premise acceptance rate and overall performance indicates that models with higher false Islamic query acceptance rates tend to exhibit lower overall reasoning accuracy. Our findings carry important implications: the path forward for Islamic NLP lies not in lightweight post-hoc techniques such as prompt or instruction tuning but in enriching knowledge. The Islamic NLP community must prioritize training models on comprehensive, large-scale Islamic legal corpora that span classical Hadith collections, jurisprudential works from the major Islamic law schools ( madhabs ), and codified legal compendia. provides the first systematic evaluation framework for Islamic legal AI systems, revealing critical limitations in platforms that Muslims increasingly rely on for spiritual guidance.

Artificial Intelligence and Law
Bond University (AU), Information Technology University (PK), Queen Mary University of London (GB), Qatar University (QA)
Peace, Justice and strong institutions
Openalex Percentile: Top 85%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.