Mobile phone model identification with a multilinear discriminant analysis approach leveraging visual features from audio recordings
摘要
The field of digital audio forensics has witnessed substantial advancements, yet it continues to face significant challenges. Mobile phone model identification can be of great importance in reconstructing evidential information in crime scene investigations and, hence, empowering law enforcement. Forensic analysts seek to determine the device used for speech recording by leveraging distinctions in internal recording sensors. However, variability introduced by factors such as speaker characteristics, speech content, and environmental conditions poses a significant challenge for accurate identification. This research proposes a robust and effective system for mobile phone model identification by incorporating advanced learning techniques, machine learning, and visual texture features. A key contribution of this work is the optimal method for converting audio signals into images using Mel spectrograms. Additionally, this work introduces a novel discriminative visual texture descriptor, the Local Phase Quantization with Enhanced Histogram Descriptor (LPQ-EHD), which captures essential texture information from the Mel spectrogram images. To further enhance class discrimination, a high-order tensor representation is employed through a multilinear subspace projection technique known as Tensor Exponential Discriminant Analysis (TEDA). This approach, applied to Multiscale LPQ-EHD features, effectively encapsulates complex and rich texture information across multiple scales. Experimental evaluations conducted on two datasets demonstrate the superiority of the proposed method over state-of-the-art approaches. Notably, identification accuracies of 99.60% and 90.12% were achieved on the MOBIPHONE and Controlled-Conditions (CC) datasets, respectively.