<p>Discourse parsing is an essential task of Natural Language Processing that aims to generate a tree-formed representation for text organization. Previous studies mainly adopted supervised methods and focused on general texts. Little attention has been paid to discourse parsing of specialized domains. This study explores the discourse parsing of Chinese government documents (CGDs). Since the discourse functional units of CGD may span multiple paragraphs and sentences, the longer documents also lead to deeper hierarchies and more branches, which challenge parser performance. To improve the discourse structure prediction of CGDs, we propose a novel CGD discourse parsing model based on reinforcement learning. Specifically, under the guidance of transition-based parsing, we first formalized the CGD discourse parsing as an action sequence prediction problem by distinguishing structural and label actions for tree generation. Based on it, we further clarified the essential components of the parsing task in a well-defined Markov decision process. Next, we constructed a transformer-based model by combining transition-based and policy-based parsing. The former offers higher computational efficiency in longer documents, while the latter improves the parse quality of CGDs by learning globally informed action selection policies. Compared with traditional transition-based methods, this framework alleviates the exposure bias and loss mismatch. The experimental results show that this model has advantages when text length exceeds 5000 words. It also achieves the highest F1 score in all semantic relation prediction tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving the discourse parsing of Chinese government documents: a policy-based reinforcement learning method

  • Xiaoyu Wang,
  • Hong Zhao,
  • Weichong Zhang,
  • Fang Wang

摘要

Discourse parsing is an essential task of Natural Language Processing that aims to generate a tree-formed representation for text organization. Previous studies mainly adopted supervised methods and focused on general texts. Little attention has been paid to discourse parsing of specialized domains. This study explores the discourse parsing of Chinese government documents (CGDs). Since the discourse functional units of CGD may span multiple paragraphs and sentences, the longer documents also lead to deeper hierarchies and more branches, which challenge parser performance. To improve the discourse structure prediction of CGDs, we propose a novel CGD discourse parsing model based on reinforcement learning. Specifically, under the guidance of transition-based parsing, we first formalized the CGD discourse parsing as an action sequence prediction problem by distinguishing structural and label actions for tree generation. Based on it, we further clarified the essential components of the parsing task in a well-defined Markov decision process. Next, we constructed a transformer-based model by combining transition-based and policy-based parsing. The former offers higher computational efficiency in longer documents, while the latter improves the parse quality of CGDs by learning globally informed action selection policies. Compared with traditional transition-based methods, this framework alleviates the exposure bias and loss mismatch. The experimental results show that this model has advantages when text length exceeds 5000 words. It also achieves the highest F1 score in all semantic relation prediction tasks.