Anjali Sharma, Mayank Aggarwal · Journal of Graphic Era University 2026 · 2026
DOI: 10.13052/jgeu0975-1416.14210
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This work presents a unified deep learning framework for automatic Hindi image caption generation in low-resource settings. The framework combines a ConvNeXt encoder, an SE-inspired adaptive attention module, and a Transformer decoder to improve visual representation learning and generate contextually relevant Hindi descriptions. The proposed framework unifies a ConvNeXt-based hierarchical visual encoder, an adaptive attention mechanism enabling dynamic saliency modulation, and a transformer decoder optimized for high-fidelity caption synthesis via multi-head self-attention. This synergy facilitates effective handling of Hindi’s morphological richness, enabling precise cross-modal alignment and robust non-sequential dependency modeling. Empirical evaluations conducted on benchmark datasets demonstrate competitive performance. The model achieves a training accuracy of 77.80%, a validation accuracy of 77.56%, and stable optimization with losses near 1.26. Caption quality metrics shows the effectiveness of the proposed framework, attaining BLEU-1/2/3/4 means of 0.8646, 0.6282, 0.5401, and 0.4429, respectively, alongside a CIDEr score of 0.8158 and METEOR of 0.6622. Additionally, low WER (0.2535) and CER (0.2607) values, coupled with an F1-Score of 0.8313, affirm the robustness and linguistic coherence of generated captions. The study contributes a scalable paradigm for multilingual captioning and establishes methodological foundations applicable to broader multimodal research within low-resource linguistic domains.
No comments yet — start the discussion below.