We present the development and evaluation of GSIP (Gibberish Speech Impression Predictor), a bi-directional GRU neural network which serves as a human impression prediction model for speech inputs, incorporating both phonetic information and prosody matrices. The objective is to select appropriate acoustic prosody for gibberish and semantic speech. Experimental validation of the proposed system was conducted through a user study. Participants ranked the system’s performance against constant prosody and random prosody patterns and completed adapted Godspeed scale questionnaires to assess their perception of the GSIP-based prosody system and conversational agents. The experiment employed three embodied conversational agents, two screen-based avatars and a physical robot. The results suggest: 1) Gibberish Speech is not so engaging for conversation; 2) higher anthropomorphism degrees create a higher perception of intelligence when agents spoke Gibberish, but that effect did not hold when speaking English. We also found that our proposed system accurately predicts human impression, but fails to generate more engaging reactions compared to constant prosody.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GSIP: A New System for Prosody Selection for Gibberish Speech

  • Antonio Galiza Cerdeira Gonzalez,
  • Ikuo Mizuuchi,
  • Bipin Indurkhya

摘要

We present the development and evaluation of GSIP (Gibberish Speech Impression Predictor), a bi-directional GRU neural network which serves as a human impression prediction model for speech inputs, incorporating both phonetic information and prosody matrices. The objective is to select appropriate acoustic prosody for gibberish and semantic speech. Experimental validation of the proposed system was conducted through a user study. Participants ranked the system’s performance against constant prosody and random prosody patterns and completed adapted Godspeed scale questionnaires to assess their perception of the GSIP-based prosody system and conversational agents. The experiment employed three embodied conversational agents, two screen-based avatars and a physical robot. The results suggest: 1) Gibberish Speech is not so engaging for conversation; 2) higher anthropomorphism degrees create a higher perception of intelligence when agents spoke Gibberish, but that effect did not hold when speaking English. We also found that our proposed system accurately predicts human impression, but fails to generate more engaging reactions compared to constant prosody.