PAZHVAK: a Word-Level Farsi Speech Corpus by University of Hormozgan
摘要
Deep learning in speech-related tasks has made significant strides, primarily driven by innovations such as transformer architectures and increased computational power. However, progress has been uneven across languages, with a strong bias toward English due to the availability of large, high-quality datasets. In contrast, Farsi (Persian) remains underrepresented, hindered by a lack of comparable resources. To bridge this gap, we introduce the first publicly available word-level Farsi speech corpus, PAZHVAK. The dataset contains 4018 unique words spoken by 61 participants (38 male, 23 female), totaling 88,535 recordings and over 56 h of audio. The recordings were captured in diverse environments and with varying intonations. Each clip, lasting between 0.5 and 7 s, is sampled at 16 kHz in mono. This corpus provides a clean and comprehensive foundation for training and evaluating speech-related models in Farsi. It is freely accessible at PAZHVAK.