The purpose of this research is to explore the method of using the Word2Vec (word to vector) model to expand the Chinese vocabulary and optimize the language model under the visual threshold of big data. By collecting a large-scale corpus of Chinese text and preprocessing it, it used the Word2Vec model to train word vectors to realize the expansion of the Chinese vocabulary library. By optimizing the language model, Chinese text can be better understood and generated. The basic principle and implementation method of the Word2Vec model was introduced. A vocabulary library expansion method based on big data was studied, and a large amount of text data was analyzed to discover new vocabulary and add it to the existing vocabulary library. The expansion of the Chinese vocabulary library and the optimization method of the language model based on the big data perspective of the Word2Vec model have effectively improved the language relevance of natural language processing tasks, with the vocabulary similarity reaching a maximum of 98.5%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Expansion of Chinese Lexicon and Language Model Optimization in the Perspective of Big Data Using the Word2Vec Model

  • Zhong Luo

摘要

The purpose of this research is to explore the method of using the Word2Vec (word to vector) model to expand the Chinese vocabulary and optimize the language model under the visual threshold of big data. By collecting a large-scale corpus of Chinese text and preprocessing it, it used the Word2Vec model to train word vectors to realize the expansion of the Chinese vocabulary library. By optimizing the language model, Chinese text can be better understood and generated. The basic principle and implementation method of the Word2Vec model was introduced. A vocabulary library expansion method based on big data was studied, and a large amount of text data was analyzed to discover new vocabulary and add it to the existing vocabulary library. The expansion of the Chinese vocabulary library and the optimization method of the language model based on the big data perspective of the Word2Vec model have effectively improved the language relevance of natural language processing tasks, with the vocabulary similarity reaching a maximum of 98.5%.