Accessibility in Image Captioning: A Comparative Study of ViT-GPT2, BLIP and GIT-Base Models
摘要
This study compares three image-captioning models—ViT-GPT2, BLIP and GIT-Base—to evaluate their effectiveness in producing accessibility-focused captions. Quantitative evaluation, done using BLEU, METEOR and SPICE, highlighted ViT-GPT2’s strength in generating captions aligned with references while the qualitative evaluation showed BLIP’s ability to generate contextual captions. The study highlights the structure used for app implementation and future scope in fusion techniques and hybrid approaches.