<p>In recent years, numerous fall detection solutions have been investigated. Multimodal systems, which integrate data from multiple sensor types, generally achieve higher accuracy and reliability than unimodal approaches. However, they often face challenges of cost, complexity, deployment, and power consumption. This paper presents an innovative audio-visual multimodal fall detection approach designed for integration into intelligent buildings. The visual component uses YOLOv8 for lightweight and fast pose estimation, combined with a Long Short-Term Memory (LSTM) network, achieving 96.95% accuracy. The audio component employs Mel spectrogram features with a Convolutional Neural Network (CNN), reaching 93.88% accuracy. These modalities are integrated through an intermediate fusion strategy, followed by a shallow neural network with a single dense layer for classification, resulting in 99.90% fall detection accuracy. Our work not only demonstrates the benefits of fusing audio and visual modalities to improve the accuracy of unimodal systems but also addresses the challenges associated with multimodal approaches. By using non-wearable audio-visual sensors, our method inherently avoids the power consumption and deployment challenges of wearable-based systems, while ensuring cost efficiency and real-time feasibility on affordable edge devices.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Audio-visual multimodal fall detection to ensure the safety of elderly people in intelligent buildings: an innovative approach using LSTM, CNN, and a shallow neural network

  • Sid Ali Bourenane,
  • Sid Ahmed Henni

摘要

In recent years, numerous fall detection solutions have been investigated. Multimodal systems, which integrate data from multiple sensor types, generally achieve higher accuracy and reliability than unimodal approaches. However, they often face challenges of cost, complexity, deployment, and power consumption. This paper presents an innovative audio-visual multimodal fall detection approach designed for integration into intelligent buildings. The visual component uses YOLOv8 for lightweight and fast pose estimation, combined with a Long Short-Term Memory (LSTM) network, achieving 96.95% accuracy. The audio component employs Mel spectrogram features with a Convolutional Neural Network (CNN), reaching 93.88% accuracy. These modalities are integrated through an intermediate fusion strategy, followed by a shallow neural network with a single dense layer for classification, resulting in 99.90% fall detection accuracy. Our work not only demonstrates the benefits of fusing audio and visual modalities to improve the accuracy of unimodal systems but also addresses the challenges associated with multimodal approaches. By using non-wearable audio-visual sensors, our method inherently avoids the power consumption and deployment challenges of wearable-based systems, while ensuring cost efficiency and real-time feasibility on affordable edge devices.