Real-Time Image Captioning for Visually Impaired Assistance Using Fine-Tuned Inception-ResNet Transfer Learning Model
International Journal of Medical Toxicology and Forensic Medicine,
Vol. 16 (2026),
1 January 2026
,
Page 1-9
https://doi.org/10.22037/ijmtfm.v16.52616
Abstract
Background: Visually challenged and blind people frequently face socio-economic barriers that can make it difficult for them to live independently and contribute fully to society. Currently, studies on assistive technology have advanced, leading to the development of a visual replacement for individuals with visual impairment. Image captioning is a task at the border between computer vision (CV) and natural language processing (NLP) that addresses making a textual explanation of the image. This paper presents a Metaheuristic Optimization-Based Image Captioning for Assisting Visually Challenged People Using Deep Learning (MOIC-AVCPDL) model to deliver image captioning tasks using advanced techniques.
Methods: The proposed MOIC-AVCPDL model employs InceptionResNetV2 for effective feature extraction, yielding rich, discriminative feature representations from the input data. The Bidirectional Gated Recurrent Unit technique is used for caption prediction. Furthermore, to maximize performance, the Multi-Objective Arithmetic Optimization Algorithm is utilized for hyperparameter tuning, enabling optimal parameter selection for improved accuracy.
Results: The proposed MOIC-AVCPDL achieved BLEU-1 scores of 78.08% on Flickr8K and 77.47% on Flickr30K, indicating improved image-captioning performance.
Conclusion: The proposed method delivers effective image captioning and has the potential to assist visually impaired people.
- Image captioning, Visually challenged people, Metaheuristic optimization, Feature extraction, Image pre-processing
How to Cite
References
[1] Faurina R, Jelita A, Vatresia A, Agustian I. Image captioning to aid blind and visually impaired outdoor navigation. Int J Artif Intell. 2023;12(3):1104-17. [DOI: 10.11591/ijai.v12.i3.pp1104-1117]
[2] Ahsan H, Bhalla N, Bhatt D, Shah K. Multi-modal image captioning for the visually impaired. arXiv [Preprint]. 2021. [DOI: 10.18653/v1/2021.naacl-srw.8]
[3] Ayadi M, Masmoudi N, Almuqren L, Alshahrani HS, Aljohani RO. Designing a novel CNN-LSTM-based model for Arabic handwritten character recognition for the visually impaired person. J Disabil Res. 2025;4(1):20240080. [DOI: 10.57197/JDR-2024-0080]
[4] Safiya KM, Pandian R. A real-time image captioning framework using computer vision to help the visually impaired. Multimed Tools Appl. 2024;83(20):59413-38. [DOI: 10.1007/s11042-023-17849-7]
[5] Valipoor M, de Antonio A, Cabrera J. Analysis and design framework for the development of indoor scene understanding assistive solutions for the person with visual impairment/blindness. Multimed Syst. 2024;30(3):152. [DOI: 10.1007/s00530-024-01350-8]
[6] Hilal AM, Alrowais F, Al-Wesabi FN, Marzouk R. Red Deer Optimization with artificial intelligence enabled image captioning system for visually impaired people. Comput Syst Sci Eng. 2023;46(2). [DOI: 10.32604/csse.2023.035529]
[7] Ayadi M, Masmoudi N, Almuqren L, Aljohani RO, Alshahrani HS. Empowering accessibility in handwritten Arabic text recognition for visually impaired individuals through optimized generative adversarial network (GAN) model. J Disabil Res. 2025;4(1):20240110. [DOI: 10.57197/JDR-2024-0110]
[8] Gupta P, Katal N. Deep learning based automatic image caption generation for visually impaired people. In: Intelligent Systems and Applications in Computer Vision. Boca Raton: CRC Press; 2023. p. 141-57. [DOI: 10.1201/9781003453406-12]
[9] Alruily M, Abd El-Aziz AA, Mostafa AM, Ezz M, Mostafa E, Alsayat A, et al. Ensemble deep learning for Alzheimer’s disease diagnosis using MRI: integrating features from VGG16, MobileNet, and InceptionResNetV2 models. PLoS One. 2025;20(4):e0318620. [DOI: 10.1371/journal.pone.0318620]
[10] Mahdi E, Barreiro CM, Cabezas X. A novel hybrid approach using an attention-based Transformer + GRU model for predicting cryptocurrency prices. 2025. [DOI: 10.3390/math13091484]
[11] Chen L, Lin X, Ma L, Wang C. A BiLSTM model enhanced with multi-objective arithmetic optimization for COVID-19 diagnosis from CT images. Sci Rep. 2025;15(1):10841. [DOI: 10.1038/s41598-025-94654-2]
[12] Adityajn105. Flickr8k dataset [Internet]. Kaggle. [Link]
[13] Hsankesara. Flickr image dataset [Internet]. Kaggle. [Link]
[14] Wang B, Wang C, Zhang Q, Su Y, Wang Y, Xu Y. Cross-lingual image caption generation based on visual attention model. IEEE Access. 2020;8:104543-54. [DOI: 10.1109/ACCESS.2020.2999568]
[15] Arasi MA, Alshahrani HM, Alruwais N, Motwakel A, Ahmed NA, Mohamed A. Automated image captioning using sparrow search algorithm with improved deep learning model. IEEE Access. 2023;11:104633-42. [DOI: 10.1109/ACCESS.2023.3317276]
- Abstract Viewed: 4 times
- PDF Downloaded: 6 times