Voice-enhanced image caption generation with vision transformers and Google Cloud TTS

Abstract Today, when it comes to the digital world, engaging image captions are crucial to such industries as travel, where the strong visuals play a crucial role in attracting the audience. Nonetheless, manual captioning is both time-consuming and usually lacks the subtlety of a scene. The paper presents an AI-based model of automatic image captioning with contextual enrichment, based on deep learning models, which have been trained using the Flickr8k dataset. We combine a pretrained Vision Transformer (ViT) that retrieves granular visual features and global contextual features of images using patch-based self-attention mechanisms with a Bidirectional Long Short-Term Memory (Bi-LSTM) model that creates coherent and human-like textual descriptions by modeling bidirectional sequence dependencies. The system is designed to improve user interaction and includes a voice-augmentation option fueled by Google Cloud Text-to-Speech to audio deliver captions written in MP3 format, and an editing interface to refine outputs, which allows flexibility, personalization, and human-in-the-loop feedback. The client-server architecture includes React and Flask as the front and back ends, respectively, which enable uploading, processing, and playback of images without any issues. The test set experimental results show better results than a CNN+LSTM baseline, with ViT+BiLSTM model achieving BLEU-1: 0.4871, BLEU-2: 0.3123, BLEU-3: 0.1899, BLEU-4: 0.1060, METEOR: 0.2694, and CL Dynamics of training indicate a higher convergence rate with best validation loss of 3.5939 at epoch 8, as compared to 3.6700 at VGG16+LSTM baseline, which highlights increased efficiency, generalization, and the possibility of transforming automated content generation in the travel and beyond sector.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-06
DOI
https://doi.org/10.1038/s41598-026-71854-y
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Voice-enhanced image caption generation with vision transformers and Google Cloud TTS

M. Venkata Krishna Reddy, Srujana Inturi, Vanitha Guda
Scientific Reports
Multimodal Machine Learning Applications
article

Voice-enhanced image caption generation with vision transformers and Google Cloud TTS

M. Venkata Krishna Reddy, Srujana Inturi, Vanitha Guda
article en

Abstract

Abstract Today, when it comes to the digital world, engaging image captions are crucial to such industries as travel, where the strong visuals play a crucial role in attracting the audience. Nonetheless, manual captioning is both time-consuming and usually lacks the subtlety of a scene. The paper presents an AI-based model of automatic image captioning with contextual enrichment, based on deep learning models, which have been trained using the Flickr8k dataset. We combine a pretrained Vision Transformer (ViT) that retrieves granular visual features and global contextual features of images using patch-based self-attention mechanisms with a Bidirectional Long Short-Term Memory (Bi-LSTM) model that creates coherent and human-like textual descriptions by modeling bidirectional sequence dependencies. The system is designed to improve user interaction and includes a voice-augmentation option fueled by Google Cloud Text-to-Speech to audio deliver captions written in MP3 format, and an editing interface to refine outputs, which allows flexibility, personalization, and human-in-the-loop feedback. The client-server architecture includes React and Flask as the front and back ends, respectively, which enable uploading, processing, and playback of images without any issues. The test set experimental results show better results than a CNN+LSTM baseline, with ViT+BiLSTM model achieving BLEU-1: 0.4871, BLEU-2: 0.3123, BLEU-3: 0.1899, BLEU-4: 0.1060, METEOR: 0.2694, and CL Dynamics of training indicate a higher convergence rate with best validation loss of 3.5939 at epoch 8, as compared to 3.6700 at VGG16+LSTM baseline, which highlights increased efficiency, generalization, and the possibility of transforming automated content generation in the travel and beyond sector.

Scientific Reports
Chaitanya Bharathi Institute of Technology (IN)
Openalex Percentile: Top 15%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.