Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

{"Arabic":[0,9,17,39,66,141,187],"large-language-model":[1],"(LLM)":[2],"evaluation":[3,122],"has":[4],"matured":[5],"around":[6],"Modern":[7],"Standard":[8],"(MSA):":[10],"aggregated":[11],"leaderboards":[12],"such":[13],"as":[14,182,212],"the":[15,41,68,158,195,199,203,248,253],"Open":[16],"LLM":[18],"Leaderboard":[19],"(OALL),":[20],"HELM":[21],"Arabic,":[22],"and":[23,32,67,97,113,139,146,168,238,247],"BALSAM":[24],"rank":[25],"models":[26,129,188],"across":[27,81,136],"dozens":[28],"of":[29,123,202,229],"MSA":[30,73,154,183,221],"tasks,":[31],"frontier":[33,128],"systems":[34,125,148],"increasingly":[35],"saturate":[36],"them.":[37],"Dialectal":[38],"-":[40,46,84,99,126,149,215],"language":[42],"Iraqis":[43],"actually":[44],"speak":[45],"remains":[47],"nearly":[48],"invisible":[49,219],"to":[50,177,220],"this":[51,244],"infrastructure.":[52],"We":[53],"introduce":[54],"Mizan":[55],"(\\"the":[56],"balance\\"),":[57],"Iraq's":[58],"national":[59],"benchmark":[60],"for":[61],"evaluating":[62],"LLMs":[63],"on":[64,116,194],"Iraqi":[65,69,79,159,196],"civic":[70],"context:":[71],"an":[72,78,140,216,226],"baseline":[74],"track":[75,80,155,160],"paired":[76],"with":[77,108,162],"six":[82],"axes":[83],"dialect":[85,87],"comprehension,":[86],"generation,":[88],"bidirectional":[89],"MSA-Iraqi":[90],"translation,":[91],"Iraq-specific":[92],"knowledge,":[93],"official-document":[94],"field":[95],"extraction,":[96],"safety":[98],"built":[100],"from":[101],"340":[102],"originally":[103],"authored,":[104],"dually":[105],"reviewed":[106],"items":[107,211],"statistically":[109,169],"audited":[110],"answer":[111],"positions":[112],"Wilson":[114],"intervals":[115],"every":[117,175],"published":[118],"score.":[119],"A":[120],"pilot":[121],"27":[124],"closed":[127],"three":[130],"days":[131],"after":[132],"release,":[133],"open":[134],"weights":[135],"size":[137],"tiers,":[138],"trio":[142],"spanning":[143],"commercial,":[144],"open-specialized,":[145],"sovereign":[147],"yields":[150],"four":[151],"findings.":[152],"The":[153,223],"saturates":[156],"while":[157],"discriminates,":[161],"a":[163,191,234],"consistent":[164],"14-18-point":[165],"per-model":[166],"gap":[167],"tied":[170],"leaders.":[171],"Official-document":[172],"extraction":[173],"confines":[174],"system":[176],"32-56.":[178],"Arabic-focused":[179],"specialization":[180],"behaves":[181],"specialization:":[184],"two":[185],"dedicated":[186],"score":[189],"below":[190],"size-matched":[192],"generalist":[193],"track.":[197],"And":[198],"safety-hardened":[200],"tier":[201],"newest":[204],"model":[205],"family":[206],"deterministically":[207],"refuses":[208],"innocuous":[209],"dialect-comprehension":[210],"policy":[213],"violations":[214],"over-refusal":[217],"mode":[218],"benchmarks.":[222],"platform":[224],"enforces":[225],"integrity":[227],"protocol":[228],"immutable":[230],"snapshots,":[231],"verification":[232],"certificates,":[233],"human":[235],"publication":[236],"gate,":[237],"public":[239,249],"retraction,":[240],"all":[241],"exercised":[242],"during":[243],"study.":[245],"Code":[246],"development":[250],"set":[251],"accompany":[252],"paper.":[254]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-11
DOI
https://doi.org/10.5281/zenodo.22714865
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

Mustafa S. Aljumaily, Nawar S. Alseelawi
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the Iraqi Civic Context

Mustafa S. Aljumaily, Nawar S. Alseelawi
preprint en

Abstract

Arabic large-language-model (LLM) evaluation has matured around Modern Standard Arabic (MSA): aggregated leaderboards such as the Open Arabic LLM Leaderboard (OALL), HELM Arabic, and BALSAM rank models across dozens of MSA tasks, and frontier systems increasingly saturate them. Dialectal Arabic - the language Iraqis actually speak - remains nearly invisible to this infrastructure. We introduce Mizan ("the balance"), Iraq's national benchmark for evaluating LLMs on Iraqi Arabic and the Iraqi civic context: an MSA baseline track paired with an Iraqi track across six axes - dialect comprehension, dialect generation, bidirectional MSA-Iraqi translation, Iraq-specific knowledge, official-document field extraction, and safety - built from 340 originally authored, dually reviewed items with statistically audited answer positions and Wilson intervals on every published score. A pilot evaluation of 27 systems - closed frontier models three days after release, open weights across size tiers, and an Arabic trio spanning commercial, open-specialized, and sovereign systems - yields four findings. The MSA track saturates while the Iraqi track discriminates, with a consistent 14-18-point per-model gap and statistically tied leaders. Official-document extraction confines every system to 32-56. Arabic-focused specialization behaves as MSA specialization: two dedicated Arabic models score below a size-matched generalist on the Iraqi track. And the safety-hardened tier of the newest model family deterministically refuses innocuous dialect-comprehension items as policy violations - an over-refusal mode invisible to MSA benchmarks. The platform enforces an integrity protocol of immutable snapshots, verification certificates, a human publication gate, and public retraction, all exercised during this study. Code and the public development set accompany the paper.

Zenodo (CERN European Organization for Nuclear Research)
University of Misan (IQ)
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.