M Gowri Shankar, D. Surendran · International Journal of Computational Intelligence Systems 2026 · 2026
DOI: 10.1007/s44196-026-01578-4
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Although video captioning has advanced significantly, it is still difficult to produce precise, context-aware, and coherent captions for a variety of complex and varied video content. Traditional approaches frequently have problems detecting and characterizing dynamic scenes, efficiently managing sequential visual data, and making sure the captions are both contextually relevant and grammatically accurate. Hence, a new model for sequential visual encoding called Global Enhanced Transformer-Tangent Dung Beetle Optimization (GET-TDBO) is introduced in this work. The input video is first obtained from the dataset, and then scene identification is done employing a Bilateral Segmentation Network (BiSeNet V2). The identified scene is then fed into the GET to produce high-quality captions. Here, the hyperparameters of GET are tuned using the proposed TDBO to enhance the quality of the generated captions. Lastly, the GPT-NeoX Large Language Model (LLM) is used to generate the caption from the sequenced words. Furthermore, GET-TDBO is examined using measures like the Metric for Evaluation of Translation with Explicit Ordering (METEOR), Mean Average Precision (mAP), Bilingual Evaluation Understudy (BLEU), Consensus-based Image Description Evaluation (CIDEr), Recall-Oriented Understudy for Gisting Evaluation-L (ROUGE-L) and Semantic Propositional Image Caption Evaluation (SPICE) and the proposed GET-TDBO attained superior values of 88.720%, 47.387%, 85.384%, 140.63, 84.956%, and 36.200%, respectively.
No comments yet — start the discussion below.