Cities worldwide create soundscapes of a broad spectrum of sounds produced by anthropogenic and natural ecosystem activities. Thus, urban sound experiences can be diverse and affect the health and well-being of urban areas in different ways. In this scenario, machine learning techniques, particularly convolutional neural networks (CNN), offer an efficient solution for automatically identifying and analyzing sounds within urban environments. This study aims to develop a CNN for the automatic classification of the soundscape taxonomy in the historic city centers of Popayán (Colombia), Santiago de Compostela (Spain), and Venice (Italy). To accomplish this, stereophonic recordings were collected using sampling grids in each city, and Mel spectrograms were employed to train the Urbanphony-CNN. The model achieved an overall F0.75-score of 0.77, precision of 0.82, and recall of 0.69 in classifying various anthropophonic and ecotopophonic sounds, underscoring the potential of artificial intelligence in urban soundscape assessment, even when the information used for model training is derived from real-world acoustic environments. Finally, since automatic classification of the urban soundscape is still an open issue, future studies should explore deeper models that consider additional aspects, such as adaptability to different cities, and include other urban soundscape variables (emotional, cultural, and/or context influences).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A CNN-Based Approach for Classifying Urban Soundscape Taxonomy in Historic Cities

  • Carlos Duran,
  • Carlos Realpe,
  • Juan Torres,
  • Julián Grijalba,
  • David Arango

摘要

Cities worldwide create soundscapes of a broad spectrum of sounds produced by anthropogenic and natural ecosystem activities. Thus, urban sound experiences can be diverse and affect the health and well-being of urban areas in different ways. In this scenario, machine learning techniques, particularly convolutional neural networks (CNN), offer an efficient solution for automatically identifying and analyzing sounds within urban environments. This study aims to develop a CNN for the automatic classification of the soundscape taxonomy in the historic city centers of Popayán (Colombia), Santiago de Compostela (Spain), and Venice (Italy). To accomplish this, stereophonic recordings were collected using sampling grids in each city, and Mel spectrograms were employed to train the Urbanphony-CNN. The model achieved an overall F0.75-score of 0.77, precision of 0.82, and recall of 0.69 in classifying various anthropophonic and ecotopophonic sounds, underscoring the potential of artificial intelligence in urban soundscape assessment, even when the information used for model training is derived from real-world acoustic environments. Finally, since automatic classification of the urban soundscape is still an open issue, future studies should explore deeper models that consider additional aspects, such as adaptability to different cities, and include other urban soundscape variables (emotional, cultural, and/or context influences).