A review of machine learning, deep learning and large language model based harmonized system code classification techniques for import and export commodities
摘要
The Harmonized System (HS) code is a systematically created, structured, multi-purpose nomenclature used to classify import and export commodities for the purposes of facilitating trade transactions, ensuring confidentiality, maintaining statistics, and transporting goods. In this context, HS code classification methods based on traditional machine learning (ML), deep learning (DL), and large language models (LLM) have attracted the attention of researchers to address the problem of accurately classifying imported and exported goods to improve the effectiveness of preventing and controlling risks associated with customs duties. Therefore, to better understand the current panorama of HS code classification, this study introduces a systematic literature review focused on ML, DL, and LLM-based HS code classification techniques of import and export commodities from the perspective of research characteristics, language setting, HS level, performance, model training, model explanation, parameter importance, model performance, and model size, etc. Our contribution is not merely listing and analyzing different approaches, but also identifying the limitations and future research challenges of HS code classification solutions. To achieve this, a search was conducted for scientific articles between 2015 and 2025, and 29 articles were selected for final analysis and investigated. Moreover, the underlying motivation is to help inform discussions among users, developers, and policymakers about HS code classification requirements, practices, and the use of HS code classification technologies.