Benchmarking large vision language models for multilingual image caption generation in low resource settings

The problem of generating image captions for low resource, right to left languages like Urdu and Arabic has not been a well explored problem in the field of multimodal learning. There is a huge gap in multilingual visual understanding as existing captioning systems are mostly developed to cater for English language. This paper introduces the multi-lingual image caption generation system to generate semantically aligned Urdu and Arabic captioning for a single image. The proposed framework generates bilingual caption descriptions by employing a vision-language (Vis-Lng) pipeline that requires no pivot-translation step at inference; the underlying bilingual training data is itself translation-derived (Sect. 3). The ZARWA-v1 corpus introduced here were constructed by translating English source captions into Urdu and Arabic. We perform a comprehensive evaluation of different multimodal models with multilingual capabilities, such as Qwen2.5-VL, Pangea, PALO, mBLIP, PaliGemma2 and LLaVA-OV, together with three additional English-centric vision-language models (Florence-2, InstructBLIP, BLIP-2) used as baselines, using zero-shot, LoRA and QLoRA fine-tuning settings. The fine-tuning process with LoRA and QLoRA further enhances the quality of captions in both languages, with Qwen2.5-VL showing the best performance when fine-tuned with QLoRA, resulting in a BLEU-4 score of 48.5 for Urdu and 50.1 for Arabic. A set of automatic evaluation metrics such as BLEU, ROUGE-L, METEOR, chrF, BERTScore, LASER, and LaBSE, and human evaluation on 5 criteria are applied. Results show that QLoRA fine-tuning can achieve competitive or better performance than full LoRA adaptation with significantly reduced computational resources. Cross-lingual semantic consistency analysis further confirms that the generated Urdu and Arabic captions remain semantically aligned across all fine-tuned models.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-21
DOI
https://doi.org/10.1038/s41598-026-70568-5
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking large vision language models for multilingual image caption generation in low resource settings

Abdu Qaid Alameri, Javed Rashid, Muhammad Shoaib Saleem, Turke Althobaiti et al.
Scientific Reports
Multimodal Machine Learning Applications
article

Benchmarking large vision language models for multilingual image caption generation in low resource settings

Abdu Qaid Alameri, Javed Rashid, Muhammad Shoaib Saleem, Turke Althobaiti, Muzammal Hussain, Imran Khan
article en

Abstract

The problem of generating image captions for low resource, right to left languages like Urdu and Arabic has not been a well explored problem in the field of multimodal learning. There is a huge gap in multilingual visual understanding as existing captioning systems are mostly developed to cater for English language. This paper introduces the multi-lingual image caption generation system to generate semantically aligned Urdu and Arabic captioning for a single image. The proposed framework generates bilingual caption descriptions by employing a vision-language (Vis-Lng) pipeline that requires no pivot-translation step at inference; the underlying bilingual training data is itself translation-derived (Sect. 3). The ZARWA-v1 corpus introduced here were constructed by translating English source captions into Urdu and Arabic. We perform a comprehensive evaluation of different multimodal models with multilingual capabilities, such as Qwen2.5-VL, Pangea, PALO, mBLIP, PaliGemma2 and LLaVA-OV, together with three additional English-centric vision-language models (Florence-2, InstructBLIP, BLIP-2) used as baselines, using zero-shot, LoRA and QLoRA fine-tuning settings. The fine-tuning process with LoRA and QLoRA further enhances the quality of captions in both languages, with Qwen2.5-VL showing the best performance when fine-tuned with QLoRA, resulting in a BLEU-4 score of 48.5 for Urdu and 50.1 for Arabic. A set of automatic evaluation metrics such as BLEU, ROUGE-L, METEOR, chrF, BERTScore, LASER, and LaBSE, and human evaluation on 5 criteria are applied. Results show that QLoRA fine-tuning can achieve competitive or better performance than full LoRA adaptation with significantly reduced computational resources. Cross-lingual semantic consistency analysis further confirms that the generated Urdu and Arabic captions remain semantically aligned across all fine-tuned models.

Scientific Reports
Khazar University (AZ), Northern Border University (SA), University of Science and Technology (YE), International Islamic University, Islamabad (PK), University of Okara (PK)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.