Enhanced Nominal Compound Chain Extraction with Boundary and Chain Information
摘要
Nominal compound chain extraction (NCCE) aims to identify and cluster nominal compounds(NC) within documents. Existing methodologies suffer from the unsatisfying performance of nominal compound chain extraction due to the incorrect identification of nominal compound boundary and the clustering errors. In this paper, we propose a joint model for the NCCE task. For document representation, a multi-head attention approach is adopted to learn the contextual representation of a document for nominal compound chain extraction; for nominal compound recognition, boundary detection is employed to enhance the model’s ability to recognize boundaries of nominal compounds; for chain extraction, chain detection is utilized to verify the semantic coherence of the predicted nominal compound chain. Experimental results show the efficiency of the information-enhanced approach in the NCCE task. Such method can also be applied in complex scenario NC recognition tasks, e.g., domain-specific NC recognition or loanword recognition.