HARNN: experiment analysis of semi-supervised learning on image captioning using Hybrid AutoEncoder-RNN

Image captioning lies at the intersection of computer vision and natural language processing, aiming to generate coherent and contextually meaningful textual descriptions for images. This capability supports applications such as assistive technologies for the visually impaired, enhanced image retrieval, and automated content generation. Conventional captioning methods predominantly rely on fully supervised learning, demanding large volumes of annotated image–caption pairs that are costly and time-consuming to obtain. This paper proposes a Hybrid Autoencoder–RNN (HARNN) model that adopts a semi-supervised learning strategy to reduce dependency on extensive labeled data while improving caption quality. The architecture combines the Xception network for visual feature extraction, an Autoencoder for compact feature representation through dimensionality reduction, and a Recurrent Neural Network for sequential caption generation. The model is implemented in Python using TensorFlow and Keras. Performance is evaluated using standard image captioning metrics, including BLEU, METEOR, ROUGE-L, and CIDEr. Comparative analysis against baseline RNN and LSTM models demonstrates that the proposed HARNN framework produces more semantically relevant and syntactically coherent captions, validating its effectiveness and efficiency for image captioning tasks.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-09
DOI
https://doi.org/10.1038/s41598-026-70326-7
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

HARNN: experiment analysis of semi-supervised learning on image captioning using Hybrid AutoEncoder-RNN

Aishwarya D Shetty, Jyothi Shetty
Scientific Reports
Multimodal Machine Learning Applications
article

HARNN: experiment analysis of semi-supervised learning on image captioning using Hybrid AutoEncoder-RNN

Aishwarya D Shetty, Jyothi Shetty
article en

Abstract

Image captioning lies at the intersection of computer vision and natural language processing, aiming to generate coherent and contextually meaningful textual descriptions for images. This capability supports applications such as assistive technologies for the visually impaired, enhanced image retrieval, and automated content generation. Conventional captioning methods predominantly rely on fully supervised learning, demanding large volumes of annotated image–caption pairs that are costly and time-consuming to obtain. This paper proposes a Hybrid Autoencoder–RNN (HARNN) model that adopts a semi-supervised learning strategy to reduce dependency on extensive labeled data while improving caption quality. The architecture combines the Xception network for visual feature extraction, an Autoencoder for compact feature representation through dimensionality reduction, and a Recurrent Neural Network for sequential caption generation. The model is implemented in Python using TensorFlow and Keras. Performance is evaluated using standard image captioning metrics, including BLEU, METEOR, ROUGE-L, and CIDEr. Comparative analysis against baseline RNN and LSTM models demonstrates that the proposed HARNN framework produces more semantically relevant and syntactically coherent captions, validating its effectiveness and efficiency for image captioning tasks.

Scientific Reports
Nitte University (IN), Visvesvaraya Technological University (IN)
Openalex Percentile: Top 15%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

HARNN: experiment analysis of semi-supervised learning on image captioning using Hybrid AutoEncoder-RNN — Aishwarya D Shetty, Jyothi Shetty · Scientific Reports (2026) | TGRS Research Map | TGRS