VISMAID: Visual Impairment Support Through Multimodal AI-Driven Description
摘要
This paper presents VISMAID (Visual Impairment Support through Multimodal AI-driven Description), a multimodal assistive system designed to enhance accessibility for visually impaired users through efficient integration of speech recognition, vision-language processing, and text-to-speech synthesis. VISMAID leverages optimized open-source AI models capable of executing entirely on local mobile hardware, addressing critical challenges such as latency variability, privacy concerns, and ongoing operational costs associated with traditional cloud-based solutions. We conduct an extensive evaluation of different multimodal AI models, focusing on their computational efficiency, mobile compatibility, licensing constraints, and multilingual capabilities. Notably, models such as Phi-4-Multimodal and LLaVA One Vision exhibit strong potential for local deployment on contemporary mobile processors including Apple’s M1, A14 Bionic, and equivalent Qualcomm Snapdragon chips. Our tests demonstrate practical inference times under one minute for the complete pipeline on these devices, confirming the feasibility of deploying sophisticated multimodal AI locally without compromising usability. Future directions discussed involve further reducing model sizes via quantization and knowledge distillation, integrating advanced open-source text-to-speech solutions, and extending support to mid-range mobile devices to broaden accessibility. VISMAID represents a viable path toward sustainable, autonomous, and universally accessible assistive technology, fostering increased independence and improved quality of life for visually impaired individuals.