Using Generative AI to Improve Library Book Classification Accuracy: The Role of Increased Training Data
摘要
This study explores the application of generative artificial intelligence (AI) in automating the assignment of classification numbers to books acquired by libraries, focusing on the Nippon Decimal Classification (NDC) system. Using a dataset from the National Diet Library (NDL) consisting of 1,000 to 1.2 million samples, the research evaluated the impact of training data volume on classification accuracy. Fine-tuning was conducted using OpenAI’s GPT-3.5 Turbo model with default parameter settings. Results indicated that even with 1,000 training samples, the model could correctly assign NDC division (two-digit) numbers to about half of the books, achieving over 70% accuracy for NDC class (one-digit) numbers. As the number of training samples increased, accuracy improved logarithmically, reaching 70.5% with 800,000 samples. However, when the sample size reached 1.2 million, there was only a marginal improvement observed to 70.7%. This trend was consistent across NDC class, division, and section (three-digit) levels, with more pronounced improvements observed in the more detailed classification levels. This study not only highlights the potential of generative AI for library classification tasks, but also suggests that more training data may not necessarily lead to a substantial improvement in accuracy. The findings indicate that while generative AI holds potential for automating library classification, the effectiveness of this approach depends on the quality and volume of the data used for training.