ConvNeXt-Driven Image Captioning in Hindi with Adaptive Attention and Transformer Decoder
DOI:
https://doi.org/10.13052/jgeu0975-1416.14210Keywords:
Hindi image captioning, ConvNxt based image encoder, Adaptive attention mechansim, Transformer decoderAbstract
This work presents a unified deep learning framework for automatic Hindi image caption generation in low-resource settings. The framework combines a ConvNeXt encoder, an SE-inspired adaptive attention module, and a Transformer decoder to improve visual representation learning and generate contextually relevant Hindi descriptions. The proposed framework unifies a ConvNeXt-based hierarchical visual encoder, an adaptive attention mechanism enabling dynamic saliency modulation, and a transformer decoder optimized for high-fidelity caption synthesis via multi-head self-attention. This synergy facilitates effective handling of Hindi’s morphological richness, enabling precise cross-modal alignment and robust non-sequential dependency modeling. Empirical evaluations conducted on benchmark datasets demonstrate competitive performance. The model achieves a training accuracy of 77.80%, a validation accuracy of 77.56%, and stable optimization with losses near 1.26. Caption quality metrics shows the effectiveness of the proposed framework, attaining BLEU-1/2/3/4 means of 0.8646, 0.6282, 0.5401, and 0.4429, respectively, alongside a CIDEr score of 0.8158 and METEOR of 0.6622. Additionally, low WER (0.2535) and CER (0.2607) values, coupled with an F1-Score of 0.8313, affirm the robustness and linguistic coherence of generated captions. The study contributes a scalable paradigm for multilingual captioning and establishes methodological foundations applicable to broader multimodal research within low-resource linguistic domains.
Downloads
References
G. Hoxha, F. Melgani, and J. Slaghenauffi, “A new CNN–RNN framework for remote sensing image captioning,” in Proc. Mediterranean and Middle-East Geoscience and Remote Sensing Symp. (M2GARSS), Tunis, Tunisia, 2020, pp. 1–4, doi: 10.1109/M2GARSS47143.2020.9105191.
H. Wang, H. Wang, and K. Xu, “Evolutionary recurrent neural network for image captioning,” Neurocomputing, vol. 401, pp. 249–256, 2020, doi: 10.1016/j.neucom.2020.03.087.
S. Jaiswal, H. Pallthadka, R. P. Chinchewadi, and T. Jaiswal, “A deep learning model for automatic image captioning using GRU and attention mechanism,” Int. J. Comput. Eng. Res. Trends, vol. 11, no. 1, pp. 28–36, Jan. 2024, doi: 10.22362/ijcert.v11i1.919.
P. Singh, C. Kumar, and A. Kumar, “Next-LSTM: A novel LSTM-based image captioning technique,” Int. J. Syst. Assur. Eng. Manag., vol. 14, pp. 1492–1503, 2023, doi: 10.1007/s13198-023-01956-7.
A. Jamil et al., “Deep learning approaches for image captioning: Opportunities, challenges and future potential,” IEEE Access, 2024, doi: 10.1109/ACCESS.2024.3365528.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Jun. 2015.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proc. 32nd Int. Conf. Mach. Learn. (ICML), Lille, France, Jul. 2015, pp. 2048–2057.
M.-T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” arXiv preprint arXiv:1508.04025, 2015.
P. Anderson, S. Gould, and M. Johnson, “Partially-supervised image captioning”, Advances in Neural Information Processing Systems, vol. 31, 2018.
A. Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale”, arXiv preprint arXiv:2010. 11929, 2020.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019.
X. Li et al., “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in Proc. Eur. Conf. Comput. Vis. (ECCV), vol. 12375, Lecture Notes in Computer Science, Springer, Cham, 2020, doi: 10.1007/978-3-030-58577-8_8.
Y.-C. Chen et al., “UNITER: Universal image–text representation learning,” in Proc. Eur. Conf. Comput. Vis. (ECCV), vol. 12375, Lecture Notes in Computer Science, Springer, Cham, 2020, doi: 10.1007/978-3-030-58577-8_7.
S. Liu, S. Cao, X. Lu, J. Peng, L. Ping, X. Fan, F. Teng, and X. Liu, “Lightweight deep learning model, ConvNeXt-U: An improved U-Net network for extracting cropland in complex landscapes from Gaofen-2 images,” Sensors, vol. 25, no. 1, Art. no. 261, 2025, doi: 10.3390/s25010261.
Z. Zhou, Y. Yang, Z. Li, et al., “Image captioning with residual Swin Transformer and actor–critic,” Neural Comput. Appl., vol. 37, pp. 8019–8031, 2025, doi: 10.1007/s00521-022-07848-4.
Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer, “Multilingual denoising pre-training for neural machine translation,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 726–742, 2020, doi: 10.1162/tacla00343.
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation,” in Proc. Int. Conf. Mach. Learn. (ICML), 2022.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” in Proc. Int. Conf. Mach. Learn. (ICML), 2023.
J. Wang, Z. Yang, X. Liu, Z. Wang, and L. Wang, “GIT: A Generative Image-to-Text Transformer for Vision and Language,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022.
T. Xian, Z. Li, C. Zhang, and H. Ma, “Dual global enhanced transformer for image captioning,” Neural Networks, vol. 148, pp. 129–141, 2022, doi: 10.1016/j.neunet.2022.01.011.
Y. Shi, J. Xia, M. Zhou, and Z. Cao, “A dual-feature-based adaptive shared transformer network for image captioning,” IEEE Trans. Instrum. Meas., vol. 73, Art. no. 5009613, pp. 1–13, 2024, doi: 10.1109/TIM.2024.3353830.
T.-Y. Lin et al., “Microsoft COCO: Common objects in context,” in Proc. Eur. Conf. Comput. Vis. (ECCV), Lecture Notes in Computer Science, vol. 8693, Springer, Cham, 2014, doi: 10.1007/978-3-319-10602-1_48.
A. Mao, M. Mohri, and Y. Zhong, “Cross-entropy loss functions: Theoretical analysis and applications,” in Proc. 40th Int. Conf. Mach. Learn. (ICML), vol. 202, Proc. Mach. Learn. Res., Jul. 2023, pp. 23803–23828.
R. Llugsi, S. E. Yacoubi, A. Fontaine, and P. Lupera, “Comparison between Adam, AdaMax and AdamW optimizers to implement a weather forecast based on neural networks for the Andean city of Quito,” in Proc. IEEE Fifth Ecuador Tech. Chapters Meeting (ETCM), Cuenca, Ecuador, 2021, pp. 1–6, doi: 10.1109/ETCM53643.2021.9590681.
M. Ghassemiazghandi, “An evaluation of ChatGPT’s translation accuracy using BLEU score,” Theory Pract. Lang. Stud., vol. 14, no. 4, pp. 985–994, Apr. 2024, doi: 10.17507/tpls.1404.07.
G. Oliveira dos Santos, E. L. Colombini, and S. Avila, “CIDEr-R: Robust consensus-based image description evaluation,” in Proc. 7th Workshop Noisy User-generated Text (W-NUT), Nov. 2021, pp. 351–360, doi: 10.18653/v1/2021.wnut-1.39.
H. Saadany and C. Orasan, “BLEU, METEOR, BERTScore: Evaluation of metrics performance in assessing critical translation errors in sentiment-oriented text,” arXiv preprint arXiv:2109.14250, 2021.


