<p>Sensor-based Human Activity Recognition (HAR) is a fundamental component of health monitoring, smart environments, and wearable intelligence; however, temporal segmentation parameters are often selected heuristically, limiting reproducibility and performance. This work proposes a systematic, data-driven framework to quantify the impact of window size (1–10&#xa0;s) and overlap ratio (25%, 50%, 75%)&#xa0;on HAR performance. A uniform segmentation strategy&#xa0;is applied across LSTM, Bi-LSTM, and CNN–BiLSTM&#xa0;architectures to ensure fair and unbiased comparison. Experiments are conducted on the&#xa0;PAMAP2 benchmark dataset, which comprises diverse daily, household, and dynamic physical activities recorded using multiple body-worn inertial sensors. Bayesian Optimization, implemented via&#xa0;Optuna with the Tree-structured Parzen Estimator (TPE), is employed to jointly optimize temporal segmentation, architectural, and training hyperparameters. The search space includes window size, overlap ratio, recurrent and convolutional units, kernel and pooling dimensions, dense-layer width, dropout rate, learning rate, batch size, optimizer selection, and training epochs. Early stopping with a patience of five epochs&#xa0;is used to ensure stable convergence. Performance is evaluated using&#xa0;Accuracy, Precision, Recall, F1-score, and Specificity&#xa0;under&#xa0;fivefold and tenfold window-level cross-validation. Results demonstrate that&#xa0;small window sizes combined with high overlap significantly enhance sensitivity to activity transitions, as confirmed through confusion-matrix-based class-wise analysis, while&#xa0;larger windows favor sustained activities with smoother temporal dynamics. The optimized&#xa0;CNN–BiLSTM&#xa0;model, using a 1&#xa0;s window with 75% overlap, achieves up to&#xa0;99.65–99.80% accuracy&#xa0;and&#xa0;99.99% specificity, outperforming recent PAMAP2 baselines by&#xa0;up to 8.35%.These findings establish&#xa0;segmentation-aware design principles&#xa0;for robust and reproducible HAR and demonstrate that&#xa0;high recognition accuracy can be achieved with low-latency inference, supporting real-time deployment on wearable and edge platforms. Future work will extend this framework to&#xa0;leave-one-subject-out (LOSO) validation&#xa0;and&#xa0;multi-dataset evaluation&#xa0;to further assess subject-independent generalization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the impact of window size, overlap, and hyperparameter tuning on the performance of deep learning models in sensor-based human activity recognition

  • Mariam El Ghazi,
  • Noura Aknin

摘要

Sensor-based Human Activity Recognition (HAR) is a fundamental component of health monitoring, smart environments, and wearable intelligence; however, temporal segmentation parameters are often selected heuristically, limiting reproducibility and performance. This work proposes a systematic, data-driven framework to quantify the impact of window size (1–10 s) and overlap ratio (25%, 50%, 75%) on HAR performance. A uniform segmentation strategy is applied across LSTM, Bi-LSTM, and CNN–BiLSTM architectures to ensure fair and unbiased comparison. Experiments are conducted on the PAMAP2 benchmark dataset, which comprises diverse daily, household, and dynamic physical activities recorded using multiple body-worn inertial sensors. Bayesian Optimization, implemented via Optuna with the Tree-structured Parzen Estimator (TPE), is employed to jointly optimize temporal segmentation, architectural, and training hyperparameters. The search space includes window size, overlap ratio, recurrent and convolutional units, kernel and pooling dimensions, dense-layer width, dropout rate, learning rate, batch size, optimizer selection, and training epochs. Early stopping with a patience of five epochs is used to ensure stable convergence. Performance is evaluated using Accuracy, Precision, Recall, F1-score, and Specificity under fivefold and tenfold window-level cross-validation. Results demonstrate that small window sizes combined with high overlap significantly enhance sensitivity to activity transitions, as confirmed through confusion-matrix-based class-wise analysis, while larger windows favor sustained activities with smoother temporal dynamics. The optimized CNN–BiLSTM model, using a 1 s window with 75% overlap, achieves up to 99.65–99.80% accuracy and 99.99% specificity, outperforming recent PAMAP2 baselines by up to 8.35%.These findings establish segmentation-aware design principles for robust and reproducible HAR and demonstrate that high recognition accuracy can be achieved with low-latency inference, supporting real-time deployment on wearable and edge platforms. Future work will extend this framework to leave-one-subject-out (LOSO) validation and multi-dataset evaluation to further assess subject-independent generalization.