Benchmarking large vision language models for multilingual image caption generation in low resource settings
The problem of generating image captions for low resource, right to left languages like Urdu and Arabic has not been a well explored problem in the field of multimodal learning. There is a huge gap in multilingual visual understanding as existing captioning systems are mostly developed to cater for English language. This paper introduces the multi-lingual image caption generation system to generate semantically aligned Urdu and Arabic captioning for a single image. The proposed framework generates bilingual caption descriptions by employing a vision-language (Vis-Lng) pipeline that requires no pivot-translation step at inference; the underlying bilingual training data is itself translation-derived (Sect. 3). The ZARWA-v1 corpus introduced here were constructed by translating English source captions into Urdu and Arabic. We perform a comprehensive evaluation of different multimodal models with multilingual capabilities, such as Qwen2.5-VL, Pangea, PALO, mBLIP, PaliGemma2 and LLaVA-OV, together with three additional English-centric vision-language models (Florence-2, InstructBLIP, BLIP-2) used as baselines, using zero-shot, LoRA and QLoRA fine-tuning settings. The fine-tuning process with LoRA and QLoRA further enhances the quality of captions in both languages, with Qwen2.5-VL showing the best performance when fine-tuned with QLoRA, resulting in a BLEU-4 score of 48.5 for Urdu and 50.1 for Arabic. A set of automatic evaluation metrics such as BLEU, ROUGE-L, METEOR, chrF, BERTScore, LASER, and LaBSE, and human evaluation on 5 criteria are applied. Results show that QLoRA fine-tuning can achieve competitive or better performance than full LoRA adaptation with significantly reduced computational resources. Cross-lingual semantic consistency analysis further confirms that the generated Urdu and Arabic captions remain semantically aligned across all fine-tuned models.
Authors
- Abdu Qaid Alameri (ORCID: https://orcid.org/0000-0002-9920-4892)
- Javed Rashid
- Muhammad Shoaib Saleem
- Turke Althobaiti
- Muzammal Hussain
- Imran Khan
Institutions
- Khazar University (AZ)
- Northern Border University (SA)
- University of Science and Technology (YE)
- International Islamic University, Islamabad (PK)
- University of Okara (PK)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1038/s41598-026-70568-5
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00