Image-to-Text Generation: Bridging Visual and Linguistic Worlds
摘要
This chapter delves into the field of image-to-text generation, a pivotal advancement in artificial intelligence that bridges computer vision and natural language processing. It provides a historical overview of image-to-text systems, from early optical character recognition (OCR) to sophisticated transformer-based and multimodal models capable of generating descriptive and contextually relevant text. Key applications across accessibility, healthcare, social media, and e-commerce underscore the transformative impact of image-to-text technology in enhancing user interaction and information accessibility. The chapter also discusses various challenges, including contextual understanding, cultural diversity, and computational efficiency, and reviews advanced techniques like Vision Transformers (ViTs), multimodal transformers, and hybrid models that enhance system capabilities. Emerging trends in the field highlight the potential for continued integration with real-time and edge devices, fostering inclusive and dynamic AI applications.