This chapter presents a comparison of the corpus-based typological classification of ten European Union languages obtained using parallel corpora, with one generated in a less controlled scenario, with non-parallel automatically annotated data. First, we described the specific pipeline that was created to extract and annotate multilingual data from the Arquivo.pt 2019 European Parliamentary Elections collection. Two new corpora for all EU languages were generated and made publicly and freely available: one composed of raw texts extracted from this collection and the other with syntactic annotation obtained automatically. Then, we presented an overview of different quantitative typological approaches developed for dependency parsing improvement and selected the most optimised ones to conduct our comparative analysis. Finally, we compared both scenarios using the same corpus-based strategy and showed that the classification obtained using the data provided by the Arquivo.pt dataset provides valuable linguistic information for this type of study, presenting similarities when compared to the classification based on parallel corpora. However, considering the dissimilarities observed, further analysis is required before validating this new method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robustness of Corpus-Based Typological Strategies for Dependency Parsing

  • Diego Alves,
  • Daniel Gomes

摘要

This chapter presents a comparison of the corpus-based typological classification of ten European Union languages obtained using parallel corpora, with one generated in a less controlled scenario, with non-parallel automatically annotated data. First, we described the specific pipeline that was created to extract and annotate multilingual data from the Arquivo.pt 2019 European Parliamentary Elections collection. Two new corpora for all EU languages were generated and made publicly and freely available: one composed of raw texts extracted from this collection and the other with syntactic annotation obtained automatically. Then, we presented an overview of different quantitative typological approaches developed for dependency parsing improvement and selected the most optimised ones to conduct our comparative analysis. Finally, we compared both scenarios using the same corpus-based strategy and showed that the classification obtained using the data provided by the Arquivo.pt dataset provides valuable linguistic information for this type of study, presenting similarities when compared to the classification based on parallel corpora. However, considering the dissimilarities observed, further analysis is required before validating this new method.