A Customizable Virtual Assistant Enhanced with Vision and Speech Capabilities
摘要
Recent advancements in Smart Assistants (SAs) as well as home automation have captured the attention of both researchers and consumers. Virtual Assistants (VAs) that are speech-enabled are commonly referred to as smart speakers. They offer a wide range of network-based services. In some cases, they seamlessly link to smart environments, introducing innovative and efficient user interfaces. Nevertheless, these devices present their own set of challenges and limitations. The distinctive features of this technology, particularly its hand-free operation handled by voice and implementation of voice-user interface create a unique landscape. Existing technology adoption models, however, fall short in offering a comprehensive explanation for the adoption of this innovative technology. To gain deeper insights into the motivation behind approving and using in-home voice assistants, this work shifts its focus towards examining voice interactions. Utilization of speech recognition for converting speech into text enables constructing an attractive personalized assistant. This development facilitates seamless tasks such as sending emails without manual typing, conducting Google searches without opening the browser, and accomplishing various daily activities including playing music or launching a preferred IDE, all with a single voice command. To assess its performance, the system is seamlessly integrated with home automation (open-source) environment and operated continuously for some days. During this period, users are encouraged to interact with it, and the results demonstrate increased accuracy, reliability and appeal.