In this paper, we present a comparative study of different approaches to extract disaster information based on news and texts provided by the Internet, which rely on Natural Language Processing (NLP) techniques. We evaluated these distinct approaches: the first approach uses JAPE rules, the second approach relies on ontologies due to their advantages in knowledge sharing and reuse, and the last approach uses a number of machine learning algorithms. JAPE rules are easy to create and are able to identify and extract named entities and specific linguistic patterns in simple systems. These rules are designed to provide broad identification and targeted application of disaster-related texts, but we find poor responsiveness at the level of complex systems. Ontologies were used to organize, structure and index knowledge, to improve the contextual understanding of information, and to increase accuracy and recall. Machine learning was explored using four algorithms to evaluate their ability to efficiently identify, predict and extract disaster information (location, type and damage). The performance evaluation of each approach showed that the ontology-based approach was the best and most effective overall, combining speed and accuracy in extracting relevant information. JAPE rules, although easy to construct and cover a wide range of domains, are limited in complex systems. Machine learning algorithms show good and convergent results, depending on the algorithm used: (SVM, LR, RFC and DTC), which indicates flexibility, adaptability, and effective integration and can provide better performance. In summary, this study shows that ontology-based approaches are the most effective for extracting disaster-related information. The integration of these techniques, especially by combining the strengths of different approaches and creating hybrid approaches, would provide more robust solutions for improving information extraction systems in disaster contexts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Disaster Information Extraction: Evaluation of NLP Techniques Using JAPE Rules, Ontologies and Machine Learning Approaches

  • Atmane Hadji,
  • Mohamed Khireddine Kholladi

摘要

In this paper, we present a comparative study of different approaches to extract disaster information based on news and texts provided by the Internet, which rely on Natural Language Processing (NLP) techniques. We evaluated these distinct approaches: the first approach uses JAPE rules, the second approach relies on ontologies due to their advantages in knowledge sharing and reuse, and the last approach uses a number of machine learning algorithms. JAPE rules are easy to create and are able to identify and extract named entities and specific linguistic patterns in simple systems. These rules are designed to provide broad identification and targeted application of disaster-related texts, but we find poor responsiveness at the level of complex systems. Ontologies were used to organize, structure and index knowledge, to improve the contextual understanding of information, and to increase accuracy and recall. Machine learning was explored using four algorithms to evaluate their ability to efficiently identify, predict and extract disaster information (location, type and damage). The performance evaluation of each approach showed that the ontology-based approach was the best and most effective overall, combining speed and accuracy in extracting relevant information. JAPE rules, although easy to construct and cover a wide range of domains, are limited in complex systems. Machine learning algorithms show good and convergent results, depending on the algorithm used: (SVM, LR, RFC and DTC), which indicates flexibility, adaptability, and effective integration and can provide better performance. In summary, this study shows that ontology-based approaches are the most effective for extracting disaster-related information. The integration of these techniques, especially by combining the strengths of different approaches and creating hybrid approaches, would provide more robust solutions for improving information extraction systems in disaster contexts.