This paper aims at contributing to the recent debate on the ‘health state’ of Italian language, with a specific focus on the Italian written by Italian university students. The topic has been addressed by the recently closed Univers-ITA project (funded in the PRIN 2017 call) from which the data analysed in this paper come. One of the research lines involved asking a sample of university students, from different Italian universities and different learning subjects, to write a formal text composed of at most 500 words. The texts have been annotated by experts and the number of occurrences of specific features related to orthography, lexicon, syntax, coherence and others have been recorded together with socio-demographic characteristics of the respondents. In order to identify groups of writers sharing the same writing features and to simultaneously allow for the correlation between the features themselves, we propose a new mixture model that models the counts as Poisson random variables whose parameters are generated according to a common factor latent variable model. Identifiability conditions are discussed. The model is applied to the project data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

What are the Main Features of the Italian Written by Italian University Students?

  • Silvia Dallari,
  • Nicola Grandi,
  • Angela Montanari

摘要

This paper aims at contributing to the recent debate on the ‘health state’ of Italian language, with a specific focus on the Italian written by Italian university students. The topic has been addressed by the recently closed Univers-ITA project (funded in the PRIN 2017 call) from which the data analysed in this paper come. One of the research lines involved asking a sample of university students, from different Italian universities and different learning subjects, to write a formal text composed of at most 500 words. The texts have been annotated by experts and the number of occurrences of specific features related to orthography, lexicon, syntax, coherence and others have been recorded together with socio-demographic characteristics of the respondents. In order to identify groups of writers sharing the same writing features and to simultaneously allow for the correlation between the features themselves, we propose a new mixture model that models the counts as Poisson random variables whose parameters are generated according to a common factor latent variable model. Identifiability conditions are discussed. The model is applied to the project data.