The sheer quantity and depth of social media data has opened a gateway to understand more about human behavior during certain conditions. The advent of topic modeling models has significantly helped uncover underlying hidden patterns and offered new perspectives on interpreting social phenomena. However, social media content is often brief, text-based, and unstructured in nature, presenting difficulties for data collection and analysis. In this paper, we assess four topic modeling techniques: Latent Semantic Analysis (LSA), Latent Dirichlet Allocation (LDA), BERT Transformer, and Llama3 with BERTopic. We used tweets created during hurricanes, wildfires, blizzards, and floods and implemented an all-encompassing preprocessing pipeline that included stop-word removal, word vectorization, data cleaning, and the manual elimination of unimportant phrases. The findings reveal that while LSA provides broad thematic summaries, LDA is more adept at identifying distinct topics. BERT with Hierarchical Clustering performs well in capturing the general thematic context, whereas Llama3 coupled with BERTopic excels in distinguishing theme clusters and highlights nuanced relationships between subjects. The comparison of these techniques provides insights into the strengths and weaknesses of different topic modeling approaches and their suitability for analyzing social media content during natural disasters. Furthermore, it sheds light on the efficacy of using traditional topic modeling models and generative models in conjunction with topic modeling to analyze Twitter data.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Deep Learning Techniques for Topic Modeling of Natural Disaster Tweets

  • Haardhik Kunder,
  • Varun Nair,
  • Dhruv Uberoi,
  • Khushali Deulkar,
  • Meera Narvekar

摘要

The sheer quantity and depth of social media data has opened a gateway to understand more about human behavior during certain conditions. The advent of topic modeling models has significantly helped uncover underlying hidden patterns and offered new perspectives on interpreting social phenomena. However, social media content is often brief, text-based, and unstructured in nature, presenting difficulties for data collection and analysis. In this paper, we assess four topic modeling techniques: Latent Semantic Analysis (LSA), Latent Dirichlet Allocation (LDA), BERT Transformer, and Llama3 with BERTopic. We used tweets created during hurricanes, wildfires, blizzards, and floods and implemented an all-encompassing preprocessing pipeline that included stop-word removal, word vectorization, data cleaning, and the manual elimination of unimportant phrases. The findings reveal that while LSA provides broad thematic summaries, LDA is more adept at identifying distinct topics. BERT with Hierarchical Clustering performs well in capturing the general thematic context, whereas Llama3 coupled with BERTopic excels in distinguishing theme clusters and highlights nuanced relationships between subjects. The comparison of these techniques provides insights into the strengths and weaknesses of different topic modeling approaches and their suitability for analyzing social media content during natural disasters. Furthermore, it sheds light on the efficacy of using traditional topic modeling models and generative models in conjunction with topic modeling to analyze Twitter data.