Enhancing Video Captioning with Global Transformer and Beetle Optimization Techniques

Although video captioning has advanced significantly, it is still difficult to produce precise, context-aware, and coherent captions for a variety of complex and varied video content. Traditional approaches frequently have problems detecting and characterizing dynamic scenes, efficiently managing sequential visual data, and making sure the captions are both contextually relevant and grammatically accurate. Hence, a new model for sequential visual encoding called Global Enhanced Transformer-Tangent Dung Beetle Optimization (GET-TDBO) is introduced in this work. The input video is first obtained from the dataset, and then scene identification is done employing a Bilateral Segmentation Network (BiSeNet V2). The identified scene is then fed into the GET to produce high-quality captions. Here, the hyperparameters of GET are tuned using the proposed TDBO to enhance the quality of the generated captions. Lastly, the GPT-NeoX Large Language Model (LLM) is used to generate the caption from the sequenced words. Furthermore, GET-TDBO is examined using measures like the Metric for Evaluation of Translation with Explicit Ordering (METEOR), Mean Average Precision (mAP), Bilingual Evaluation Understudy (BLEU), Consensus-based Image Description Evaluation (CIDEr), Recall-Oriented Understudy for Gisting Evaluation-L (ROUGE-L) and Semantic Propositional Image Caption Evaluation (SPICE) and the proposed GET-TDBO attained superior values of 88.720%, 47.387%, 85.384%, 140.63, 84.956%, and 36.200%, respectively.

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence Systems
Published
2026-09-18
DOI
https://doi.org/10.1007/s44196-026-01578-4
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Enhancing Video Captioning with Global Transformer and Beetle Optimization Techniques

D. Surendran, M Gowri Shankar
International Journal of Computational Intelligence Systems
Multimodal Machine Learning Applications
article

Enhancing Video Captioning with Global Transformer and Beetle Optimization Techniques

D. Surendran, M Gowri Shankar
article en

Abstract

Although video captioning has advanced significantly, it is still difficult to produce precise, context-aware, and coherent captions for a variety of complex and varied video content. Traditional approaches frequently have problems detecting and characterizing dynamic scenes, efficiently managing sequential visual data, and making sure the captions are both contextually relevant and grammatically accurate. Hence, a new model for sequential visual encoding called Global Enhanced Transformer-Tangent Dung Beetle Optimization (GET-TDBO) is introduced in this work. The input video is first obtained from the dataset, and then scene identification is done employing a Bilateral Segmentation Network (BiSeNet V2). The identified scene is then fed into the GET to produce high-quality captions. Here, the hyperparameters of GET are tuned using the proposed TDBO to enhance the quality of the generated captions. Lastly, the GPT-NeoX Large Language Model (LLM) is used to generate the caption from the sequenced words. Furthermore, GET-TDBO is examined using measures like the Metric for Evaluation of Translation with Explicit Ordering (METEOR), Mean Average Precision (mAP), Bilingual Evaluation Understudy (BLEU), Consensus-based Image Description Evaluation (CIDEr), Recall-Oriented Understudy for Gisting Evaluation-L (ROUGE-L) and Semantic Propositional Image Caption Evaluation (SPICE) and the proposed GET-TDBO attained superior values of 88.720%, 47.387%, 85.384%, 140.63, 84.956%, and 36.200%, respectively.

International Journal of Computational Intelligence Systems
Government of Tamil Nadu (IN), Karpagam Academy of Higher Education (IN)
Quality Education
Openalex Percentile: Top 13%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Enhancing Video Captioning with Global Transformer and Beetle Optimization Techniques — D. Surendran, M Gowri Shankar · International Journal of Computational Intelligence Systems (2026) | TGRS Research Map | TGRS