<p>This study introduces a comprehensive dataset of co-occurring Twitter hashtags, specifically designed to benchmark the performance of graph neural network (GNN) models in node classification tasks, with a focus on hashtag virality. The dataset, collected over 10&#xa0;days and comprising 4655 nodes and 5901 edges, incorporates a diverse set of 12 features for each hashtag. These features include standard metrics, such as structural characteristics within co-occurrence networks, and novel attributes, including semantic similarity, polarity, subjectivity, and a unique user activity measure based on activity days and tweet volume. A new metric for determining user activity levels on social networks is introduced and used as a feature in the dataset. The dataset is divided into two parts: a feature matrix representing each hashtag’s attributes and a connection matrix detailing the co-occurrence relationships between hashtags. Labels indicating viral or non-viral behavior were assigned during data collection, providing a valuable resource for evaluating the effectiveness of GNN models in classifying hashtag virality. The high F1-scores achieved using these models underscore the significance of integrating structural, content-based, and user-related features, as well as the relationships between hashtags, in accurately classifying hashtag virality on Twitter. This work also demonstrates the significant influence of the selected features and relationships within the dataset, highlighting the importance of a comprehensive graph-based approach in understanding and predicting hashtag virality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TwitterTagNet: an extensive graph dataset for node classification in co-occurring hashtag networks

  • Sina Firuzbakht,
  • Mohammad Khansari

摘要

This study introduces a comprehensive dataset of co-occurring Twitter hashtags, specifically designed to benchmark the performance of graph neural network (GNN) models in node classification tasks, with a focus on hashtag virality. The dataset, collected over 10 days and comprising 4655 nodes and 5901 edges, incorporates a diverse set of 12 features for each hashtag. These features include standard metrics, such as structural characteristics within co-occurrence networks, and novel attributes, including semantic similarity, polarity, subjectivity, and a unique user activity measure based on activity days and tweet volume. A new metric for determining user activity levels on social networks is introduced and used as a feature in the dataset. The dataset is divided into two parts: a feature matrix representing each hashtag’s attributes and a connection matrix detailing the co-occurrence relationships between hashtags. Labels indicating viral or non-viral behavior were assigned during data collection, providing a valuable resource for evaluating the effectiveness of GNN models in classifying hashtag virality. The high F1-scores achieved using these models underscore the significance of integrating structural, content-based, and user-related features, as well as the relationships between hashtags, in accurately classifying hashtag virality on Twitter. This work also demonstrates the significant influence of the selected features and relationships within the dataset, highlighting the importance of a comprehensive graph-based approach in understanding and predicting hashtag virality.