Reading Between the Lines: Detecting Corporate Financial Fraud Using Multi-dimensional Textual Features
摘要
In this study, we propose a novel resampling algorithm which incorporates multi-dimensional textual features to address the challenges in corporate financial fraud detection. Using a sample of Chinese listed firms, we apply various text mining techniques to extract similarity, readability and sentiment information of corporate annual reports measured by a collection of textual features. We improve the SMOTE (Synthetic Minority Oversampling Technique) algorithm by calculating the weighted Euclidean distance based on feature importance to generate new instances and rebalance the sample distribution. The overall predictive accuracy significantly benefits from the inclusion of textual features and the application of the proposed resampling method, which remains robust when accounting for the misclassification costs. We analyze the contributions of predictors and interpret the model at the end. Our study shows the power of soft information from corporate disclosure for fraud detection and provides new algorithms in solving business ethics questions.