A Method of Identifying Russian-Language Machine-Generated Texts Based on Information Packaging
摘要
The development of a method to detect machine-generated texts, based on regularities of information packaging in the Russian language, is considered. It is shown that the problem of detecting machine-generated texts is especially relevant for scientific texts that unscrupulous authors use when writing scientific and student papers. The capabilities of the SimAlign package for displaying word order correspondence in Russian- and English-language sentences are studied. The difference in the information packaging of Russian- and English-language texts is illustrated, and statistical data on the frequency of changes in the structures of a Russian-language sentence and its machine translation into English are provided. A method for detecting machine-generated Russian-language texts is proposed, which is based on the differences in information packaging in languages with free and fixed word order.