HARNN: experiment analysis of semi-supervised learning on image captioning using Hybrid AutoEncoder-RNN
Image captioning lies at the intersection of computer vision and natural language processing, aiming to generate coherent and contextually meaningful textual descriptions for images. This capability supports applications such as assistive technologies for the visually impaired, enhanced image retrieval, and automated content generation. Conventional captioning methods predominantly rely on fully supervised learning, demanding large volumes of annotated image–caption pairs that are costly and time-consuming to obtain. This paper proposes a Hybrid Autoencoder–RNN (HARNN) model that adopts a semi-supervised learning strategy to reduce dependency on extensive labeled data while improving caption quality. The architecture combines the Xception network for visual feature extraction, an Autoencoder for compact feature representation through dimensionality reduction, and a Recurrent Neural Network for sequential caption generation. The model is implemented in Python using TensorFlow and Keras. Performance is evaluated using standard image captioning metrics, including BLEU, METEOR, ROUGE-L, and CIDEr. Comparative analysis against baseline RNN and LSTM models demonstrates that the proposed HARNN framework produces more semantically relevant and syntactically coherent captions, validating its effectiveness and efficiency for image captioning tasks.
Authors
- Aishwarya D Shetty (ORCID: https://orcid.org/0000-0002-7302-7285)
- Jyothi Shetty
Institutions
- Nitte University (IN)
- Visvesvaraya Technological University (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1038/s41598-026-70326-7
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00