In this paper, we present an analysis of corpus homogeneity, that is, the similar distribution of (linguistic) features throughout one corpus. The research is motivated by contrastive interlanguage analysis (CIA) of the Belgisches Deutschkorpus (Beldeko), in which we compare the characteristics of cohesion in second language (L2) learner writing with first language (L1) writing. Recent research has highlighted the importance of corpus homogeneity for CIA (Granger 2021; Shadrova et al. 2021). Following a number of suggested procedures, we analyse the levels of corpus homogeneity of coarse linguistic features (part of speech, POS) and a more fine-grained linguistic feature (connectives). Two hypotheses were investigated: (1) based on previous research on L1 German, the texts should show similar patterns in the distribution of POS; (2) since the use of cohesive devices is related to individual writing style, higher heterogeneity for these elements is a likely finding. The findings confirm our hypotheses: we found relatively homogeneous patterns in the distribution of POS, whereas connectives show less homogenous distributions. Consequently, we conclude that the results are promising in terms of Beldeko being used as a representative corpus to investigate and compare (cohesion in) L2 German of writers with L1 Dutch.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Beldeko Corpus as a Resource to Investigate Cohesion in German Learner Language: A Preliminary Analysis of Corpus Homogeneity

  • Helena Wedig,
  • Carola Strobl,
  • Jim J. J. Ureel,
  • Tanja Mortelmans

摘要

In this paper, we present an analysis of corpus homogeneity, that is, the similar distribution of (linguistic) features throughout one corpus. The research is motivated by contrastive interlanguage analysis (CIA) of the Belgisches Deutschkorpus (Beldeko), in which we compare the characteristics of cohesion in second language (L2) learner writing with first language (L1) writing. Recent research has highlighted the importance of corpus homogeneity for CIA (Granger 2021; Shadrova et al. 2021). Following a number of suggested procedures, we analyse the levels of corpus homogeneity of coarse linguistic features (part of speech, POS) and a more fine-grained linguistic feature (connectives). Two hypotheses were investigated: (1) based on previous research on L1 German, the texts should show similar patterns in the distribution of POS; (2) since the use of cohesive devices is related to individual writing style, higher heterogeneity for these elements is a likely finding. The findings confirm our hypotheses: we found relatively homogeneous patterns in the distribution of POS, whereas connectives show less homogenous distributions. Consequently, we conclude that the results are promising in terms of Beldeko being used as a representative corpus to investigate and compare (cohesion in) L2 German of writers with L1 Dutch.