Attribute knowledge and KBGAT for predicting the accuracy of the harmonized system code for classifying import and export commodities
摘要
The harmonized system (HS) code is a unique identifier for classifying import and export commodities, serving as the foundation for customs statistics and taxation and the safe regulation of imports and exports. However, its accuracies in import- and export-commodity classifications can only be examined via highly specialized and labor-intensive procedures. The traditional methods for predicting the accuracy of the HS code are limited by several factors, such as disordered declaration information, isolated declaration elements, excessively specialized terminologies, and imbalanced commodity categories, and often focus on only one of the semantic or spatial features, which results in unsatisfactory prediction accuracy. Thus, this study fully combines the semantic and spatial features of commodity description information, aiming to demonstrate a method for predicting the HS codes for classifying import and export commodities by integrating attribute knowledge with a knowledge-based graph attention network (KBGAT). Additionally, this study leverages the semantic associations and attribute associations of declaration elements in the declaration information of customs declaration, the HS code prediction problem is transformed into a link complementation problem on the knowledge graph through the structured representation of the knowledge graph and the graph attention mechanism. In comparative experiments involving single- and multi-category HS-code predictions, our KBGAT model significantly outperforms the TransE, ConvE, relational graph convolutional network (R-GCN) and bidirectional encoder representations from transformers (BERT) models across various evaluation metrics, including accuracy, F1 score, Hits@3 and Hits@10. Furthermore, the results of the ablation study indicate that the attribute associations among the declaration elements significantly improve the HS-code prediction accuracy, whereas the semantic associations in the commodity-description texts positively contribute to the prediction effectiveness.