Corpus Representativeness in Data-Driven Learning
摘要
Data-driven learning (DDL) is an approach to language learning and teaching that relies on the identification of linguistic patterns in authentic language use. Therefore, corpora, which are collections of spoken or written texts, are key to DDL practices. For corpora to yield useful findings, however, they need to be representative of the target language domain. This chapter introduces the concept of corpus representativeness and discusses its current role in DDL research and practice.