This study explores the possibilities of processing multi-textual data (objects represented by a set of texts and metadata about them). During the research, a dataset of such data was collected, containing information about 13,117 video games. Each game is represented by 2 different types of texts, the quantity of which is not known in advance. When building models, the authors were guided by the following assumptions: each individual text is meaningful and complete (therefore separate processing of texts contributes to improving predictions), texts of different types are significantly different (therefore different types of texts should be processed separately), and inclusion of metadata increases the amount of information about the object (therefore contributes to improving predictions). Within the study, 2 models using classical NLP methods and 3 models using the above assumptions were created. As a result of 2 series of experiments, the proposed models showed higher results compared to classical methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Big Data Analytics Approach with Multiple Text Types: The Case of the Computer Gaming

  • Aleksandr Belov,
  • Feodor Zakharov,
  • Egor Litvinenko,
  • Ruslan Molokanov,
  • Karina Malyshkina,
  • Ilya Semichasnov,
  • Aleksey Markin

摘要

This study explores the possibilities of processing multi-textual data (objects represented by a set of texts and metadata about them). During the research, a dataset of such data was collected, containing information about 13,117 video games. Each game is represented by 2 different types of texts, the quantity of which is not known in advance. When building models, the authors were guided by the following assumptions: each individual text is meaningful and complete (therefore separate processing of texts contributes to improving predictions), texts of different types are significantly different (therefore different types of texts should be processed separately), and inclusion of metadata increases the amount of information about the object (therefore contributes to improving predictions). Within the study, 2 models using classical NLP methods and 3 models using the above assumptions were created. As a result of 2 series of experiments, the proposed models showed higher results compared to classical methods.