With the growing popularity of social media platforms, extracting valuable information from the vast amount of text and image data has become a challenging task. This paper proposes a hybrid feature selection approach for identifying optimal features in text and image data from Twitter. The methodology involves three stages: data preprocessing, feature selection, and classification. In the data preprocessing stage, various techniques such as tokenization, stemming, and image resizing are employed to preprocess the raw text and image data. This ensures that the data is standardized and ready for further analysis. The feature selection stage combines filter and wrapper methods. Filter methods identify informative features using statistical measures, while wrapper methods evaluate their impact on classification performance. In the classification stage, different algorithms like Naive Bayes, Support Vector Machine, and Convolutional Neural Networks are used to assess the selected features’ performance using metrics like accuracy, precision, recall, and F1-score. Experimental results demonstrate that the proposed hybrid feature selection approach outperforms individual filter and wrapper methods. The selected optimal features exhibit improved classification accuracy and provide valuable insights into the text and image data on Twitter. The approach shows potential for applications such as sentiment analysis, topic classification, and user profiling.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid Feature Selection Approach to Identify Optimal Features of Text and Image in Twitter

  • Sumit Jain,
  • Rashmi Yadav,
  • Swapnil Waghela,
  • Ayesha Mandloi,
  • Nitish Pathak

摘要

With the growing popularity of social media platforms, extracting valuable information from the vast amount of text and image data has become a challenging task. This paper proposes a hybrid feature selection approach for identifying optimal features in text and image data from Twitter. The methodology involves three stages: data preprocessing, feature selection, and classification. In the data preprocessing stage, various techniques such as tokenization, stemming, and image resizing are employed to preprocess the raw text and image data. This ensures that the data is standardized and ready for further analysis. The feature selection stage combines filter and wrapper methods. Filter methods identify informative features using statistical measures, while wrapper methods evaluate their impact on classification performance. In the classification stage, different algorithms like Naive Bayes, Support Vector Machine, and Convolutional Neural Networks are used to assess the selected features’ performance using metrics like accuracy, precision, recall, and F1-score. Experimental results demonstrate that the proposed hybrid feature selection approach outperforms individual filter and wrapper methods. The selected optimal features exhibit improved classification accuracy and provide valuable insights into the text and image data on Twitter. The approach shows potential for applications such as sentiment analysis, topic classification, and user profiling.