Effective public speaking requires proficiency in pronunciation, fluency, grammar, and confidence. Existing speech evaluation systems primarily focus on isolated aspects such as pronunciation or grammatical correctness but fail to provide holistic feedback. This study introduces an AI-powered speech feedback system that integrates OpenAI’s Whisper for high-accuracy speech-to-text transcription and LLaMA, executed locally using the Ollama framework, for grammatical, structural, and sentiment analysis. Whisper, a state-of-the-art model, utilizes an encoder-decoder architecture to transcribe speech accurately across diverse accents and environments. LLaMA processes this text to evaluate grammar, readability, and coherence while performing sentiment and confidence analysis. The model quantifies confidence levels based on prosodic features, speech pace, and voice modulation. The system has been tested on 17 recorded speeches, including TED Talks, to assess pronunciation clarity, speech rate, filler word frequency, and confidence levels. Experimental results show that the system effectively identifies mispronunciations, fluency issues, filler words, and structural weaknesses. Key feedback metrics include confidence scores, cohesion analysis, and personalized recommendations to improve public speaking skills. Additionally, the system supports multilingual speech evaluation and can be adapted to various speaking styles, making it a versatile tool for enhancing communication skills across diverse user groups.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AI-Powered Speech Feedback System for Public Speaking with Local Language Support and Confidence Evaluation

  • Ansh Bhardwaj,
  • Mohd Shoaib Khan,
  • Krishna Sanil,
  • Sai Yashwanth Kotagiri,
  • L. Pranav

摘要

Effective public speaking requires proficiency in pronunciation, fluency, grammar, and confidence. Existing speech evaluation systems primarily focus on isolated aspects such as pronunciation or grammatical correctness but fail to provide holistic feedback. This study introduces an AI-powered speech feedback system that integrates OpenAI’s Whisper for high-accuracy speech-to-text transcription and LLaMA, executed locally using the Ollama framework, for grammatical, structural, and sentiment analysis. Whisper, a state-of-the-art model, utilizes an encoder-decoder architecture to transcribe speech accurately across diverse accents and environments. LLaMA processes this text to evaluate grammar, readability, and coherence while performing sentiment and confidence analysis. The model quantifies confidence levels based on prosodic features, speech pace, and voice modulation. The system has been tested on 17 recorded speeches, including TED Talks, to assess pronunciation clarity, speech rate, filler word frequency, and confidence levels. Experimental results show that the system effectively identifies mispronunciations, fluency issues, filler words, and structural weaknesses. Key feedback metrics include confidence scores, cohesion analysis, and personalized recommendations to improve public speaking skills. Additionally, the system supports multilingual speech evaluation and can be adapted to various speaking styles, making it a versatile tool for enhancing communication skills across diverse user groups.