In the world of internet people are dependent on the data available in digital form, due to which millions of information is being generated every day over the internet. The busy schedule of day today’s life had made it is a mandatory requirement to have an automation to extract the useful information from the data available over the internet in a concise timeframe. The short description of information can be generated with the help of automatic text summarization. The available systems work efficiently on English language, as India is a diverse country with many languages being spoken across the country and due to which public even use their native languages to convey the information through digital platform. Kannada is one of the prominent languages used by the southern India and 51 million population of all over the world use this language for communication, very less work has been carried out on Kannada language in the domain of Natural Language Processing. The proposed model is implemented on Kannada language for automatic text summarization using extractive method based on sentence ranking algorithm and K means clustering algorithm and got better result for sentence ranking approach with F-score of 0.68 on Rouge 1, F-score of 0.4 on Rouge 2 and F-score of 0.4 on Rouge L over 10,000 documents.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing Extractive Summarization Methods for Kannada Using Sentence Ranking and K-Means Algorithm

  • Dakshayani Ijeri,
  • Pushpa B. Patil,
  • Sanjana Mirajkar

摘要

In the world of internet people are dependent on the data available in digital form, due to which millions of information is being generated every day over the internet. The busy schedule of day today’s life had made it is a mandatory requirement to have an automation to extract the useful information from the data available over the internet in a concise timeframe. The short description of information can be generated with the help of automatic text summarization. The available systems work efficiently on English language, as India is a diverse country with many languages being spoken across the country and due to which public even use their native languages to convey the information through digital platform. Kannada is one of the prominent languages used by the southern India and 51 million population of all over the world use this language for communication, very less work has been carried out on Kannada language in the domain of Natural Language Processing. The proposed model is implemented on Kannada language for automatic text summarization using extractive method based on sentence ranking algorithm and K means clustering algorithm and got better result for sentence ranking approach with F-score of 0.68 on Rouge 1, F-score of 0.4 on Rouge 2 and F-score of 0.4 on Rouge L over 10,000 documents.