Studying human modality preferences in a human-drone framework for secondary task selection
摘要
Existing research works on Unmanned Aerial Vehicles (UAVs) equipped with aerial manipulators usually rely on controllers, which often necessitate the use of both hands, unless the drone operates autonomously. In challenging scenarios such as inspecting high-voltage power lines or repairing wind turbine blades, operators must control the drone with precise manual input while simultaneously managing aerial manipulation tasks at the site. However, achieving precise control near the task site is quite demanding, and adding aerial tasks further complicates execution, heightening cognitive load and increasing the risk of failure. In such a case, the user has to bear the responsibility of understanding the spatial relationship between the drone and the target location where the aerial task needs to be performed. This paper proposes a novel human-drone interaction framework for an effective human-drone team by leveraging a computer vision pipeline and alternative modalities such as eye gaze and speech to address the challenge and provide critical support to operators. To evaluate the effectiveness of these modalities, we conducted two exploratory studies within a mixed reality environment for comparison between them. The first study focused on a simple pointing and selection task with the user hovering the drone near a fixed target. The second study is an extension of the previous study that focuses on specific task-oriented applications. Task completion time, user cognition, and evaluated system usability were measured using the NASA Task Load Index (TLX). Participants, including those with no prior experience in drone piloting or MR interfaces, could complete the aerial manipulation tasks using the proposed interactive system. The results revealed that eye gaze is better suited for simple and fast task execution whereas speech is better for tasks requiring stability and focused awareness of the drone.