Abstract <p>The development of a method to detect machine-generated texts, based on regularities of information packaging in the Russian language, is considered. It is shown that the problem of detecting machine-generated texts is especially relevant for scientific texts that unscrupulous authors use when writing scientific and student papers. The capabilities of the SimAlign package for displaying word order correspondence in Russian- and English-language sentences are studied. The difference in the information packaging of Russian- and English-language texts is illustrated, and statistical data on the frequency of changes in the structures of a Russian-language sentence and its machine translation into English are provided. A method for detecting machine-generated Russian-language texts is proposed, which is based on the differences in information packaging in languages with free and fixed word order<i>.</i></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Method of Identifying Russian-Language Machine-Generated Texts Based on Information Packaging

  • Yu. I. Butenko

摘要

Abstract

The development of a method to detect machine-generated texts, based on regularities of information packaging in the Russian language, is considered. It is shown that the problem of detecting machine-generated texts is especially relevant for scientific texts that unscrupulous authors use when writing scientific and student papers. The capabilities of the SimAlign package for displaying word order correspondence in Russian- and English-language sentences are studied. The difference in the information packaging of Russian- and English-language texts is illustrated, and statistical data on the frequency of changes in the structures of a Russian-language sentence and its machine translation into English are provided. A method for detecting machine-generated Russian-language texts is proposed, which is based on the differences in information packaging in languages with free and fixed word order.