Readability assessment for book-level long text is widely needed in real educational applications. However, most of the current researches focus on passage-level readability assessment and little work has been done to process ultra-long texts. In order to process the long sequence of book texts better and to enhance pretrained models with difficulty knowledge, we propose a novel model DSDR, difficulty-aware segment pre-training and difficulty multi-view representation. Specifically, we split all books into multiple fixed-length segments and employ unsupervised clustering to obtain difficulty-aware segments, which are used to re-train the pretrained model to learn difficulty knowledge. Accordingly, a long text is represented by averaging multiple vectors of segments with varying difficulty levels. We construct a new dataset of Graded Children’s Books to evaluate model performance. Our proposed model achieves promising results, outperforming both the traditional SVM classifier and several popular pretrained models. In addition, our work establishes a new prototype for book-level readability assessment, which provides an important benchmark for related research in future work.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Going Beyond Passages: Readability Assessment for Book-Level Long Texts

  • Wenbiao Li,
  • Rui Sun,
  • Tianyi Zhang,
  • Yunfang Wu

摘要

Readability assessment for book-level long text is widely needed in real educational applications. However, most of the current researches focus on passage-level readability assessment and little work has been done to process ultra-long texts. In order to process the long sequence of book texts better and to enhance pretrained models with difficulty knowledge, we propose a novel model DSDR, difficulty-aware segment pre-training and difficulty multi-view representation. Specifically, we split all books into multiple fixed-length segments and employ unsupervised clustering to obtain difficulty-aware segments, which are used to re-train the pretrained model to learn difficulty knowledge. Accordingly, a long text is represented by averaging multiple vectors of segments with varying difficulty levels. We construct a new dataset of Graded Children’s Books to evaluate model performance. Our proposed model achieves promising results, outperforming both the traditional SVM classifier and several popular pretrained models. In addition, our work establishes a new prototype for book-level readability assessment, which provides an important benchmark for related research in future work.