ArGemma : A Multi‐Task Fine‐Tuning Framework for Adapting Gemma to Arabic

ABSTRACT Open‐source large language models (LLMs) have significantly advanced natural language processing (NLP), particularly for English. However, their performance in Arabic has remained limited due to the scarcity of high‐quality datasets and the high computational cost of full fine‐tuning. To address this challenge, we curated and prepared a diverse collection of Arabic‐specific datasets covering translation, summarization, storytelling, dialogue generation, and Classical–Modern Arabic bridging. Building on these resources, we developed ArGemma, an Arabic‐adapted version of the Gemma model, fine‐tuned using Supervised Fine‐Tuning (SFT) with Low‐Rank Adaptation (LoRA) to enable efficient adaptation with reduced computational overhead. A hybrid approach integrating SFT, Retrieval‐Augmented Generation (RAG), and few‐shot prompting was also employed to enhance contextual understanding and performance—particularly for tasks requiring nuanced reasoning across different forms of Arabic. ArGemma's performance was evaluated using both quantitative metrics (BLEU, ROUGE) and qualitative assessments by native speakers and language experts to measure fluency, coherence, and contextual accuracy. Results demonstrate that LoRA‐based fine‐tuning, supported by carefully prepared datasets and retrieval‐augmented techniques, can significantly enhance Arabic NLP performance while remaining computationally efficient. This work, which was awarded third place in the 2025 Google—Unlock Global Communication with Gemma Kaggle competition, demonstrates how parameter‐efficient fine‐tuning (PEFT) combined with targeted dataset preparation can effectively expand LLM support for underrepresented languages like Arabic.

Authors

Institutions

Publication Details

Journal
Expert Systems
Published
2026-09-24
DOI
https://doi.org/10.1111/exsy.70431
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ArGemma : A Multi‐Task Fine‐Tuning Framework for Adapting Gemma to Arabic

Beyza Eken, Serap Çakar Kaman, Taha Alselwi
Expert Systems
Topic Modeling
article

ArGemma : A Multi‐Task Fine‐Tuning Framework for Adapting Gemma to Arabic

Beyza Eken, Serap Çakar Kaman, Taha Alselwi
article en

Abstract

ABSTRACT Open‐source large language models (LLMs) have significantly advanced natural language processing (NLP), particularly for English. However, their performance in Arabic has remained limited due to the scarcity of high‐quality datasets and the high computational cost of full fine‐tuning. To address this challenge, we curated and prepared a diverse collection of Arabic‐specific datasets covering translation, summarization, storytelling, dialogue generation, and Classical–Modern Arabic bridging. Building on these resources, we developed ArGemma, an Arabic‐adapted version of the Gemma model, fine‐tuned using Supervised Fine‐Tuning (SFT) with Low‐Rank Adaptation (LoRA) to enable efficient adaptation with reduced computational overhead. A hybrid approach integrating SFT, Retrieval‐Augmented Generation (RAG), and few‐shot prompting was also employed to enhance contextual understanding and performance—particularly for tasks requiring nuanced reasoning across different forms of Arabic. ArGemma's performance was evaluated using both quantitative metrics (BLEU, ROUGE) and qualitative assessments by native speakers and language experts to measure fluency, coherence, and contextual accuracy. Results demonstrate that LoRA‐based fine‐tuning, supported by carefully prepared datasets and retrieval‐augmented techniques, can significantly enhance Arabic NLP performance while remaining computationally efficient. This work, which was awarded third place in the 2025 Google—Unlock Global Communication with Gemma Kaggle competition, demonstrates how parameter‐efficient fine‐tuning (PEFT) combined with targeted dataset preparation can effectively expand LLM support for underrepresented languages like Arabic.

Expert SystemsVol. 43(11)
Sakarya University (TR)
Quality Education
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.