Mining textual fields from patent documents: systematic review
摘要
Patent databases serve as a primary source of technical intelligence, offering insights into recent and emerging technologies across various domains. Text mining plays an important role in extracting this intelligence, though the process is complicated by the vast volume of data, the complex structure of patent texts, and their distinctive characteristics—including a blend of legal and technical language, multilingual content, and semi-structured data formats. This study investigates text mining methods applied to patent documents through a systematic literature review (SLR) of research published between 2018 and 2025 in the Scopus and Web of Science databases. A total of 117 scientific articles and conference papers were analyzed, enabling the identification of key themes: (1) trends in patent text mining; (2) predominant methodologies and recommended tools for preprocessing and analysis; and (3) assessments of practical implications and limitations. As a practical and managerial contribution, this SLR outlines major methodological advancements and emerging trends in the field, synthesizing key recommendations from tested approaches to highlight future research opportunities. Analyzing the textual content of patent documents enables the extraction of technological intelligence—scientific and technical knowledge that addresses real-world applications. This intelligence can support competitive advantage and strategic decision-making, ultimately serving as a powerful tool to advance the Sustainable Development Goals (SDGs) outlined in the 2030 Agenda.