This study investigates the scaling properties and dynamics of hashtags in tweets generated by Large Language Models (LLMs) to evaluate their performance. The study reveals that hashtags exhibit two scaling properties: Zipf’s law and Taylor’s law. However, these scaling properties have distinct characteristics compared to those observed in tweets generated by humans and natural language. Note that the relationship between the number of tokens and types does not follow Heaps’ law in contrast to human-generated tweets. The study calculates the generalized Jensen-Shannon divergence between hashtag distributions obtained at different time intervals to capture the dynamics quantitatively. The analysis shows that the decay of similarity between hashtag frequency distributions over time intervals follows a sublinear pattern similar to the dynamic behavior of hashtags in tweets generated by humans but differs from the dynamic characteristics of natural language. These findings suggest that although LLMs use hashtags differently from humans, they exhibit certain similarities in statistical properties and dynamics. The study is significant in understanding LLMs’ behavior on social media platforms and can help inform future AI text-generation strategies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Large Language Models on Twitter Based on Hashtag Dynamics and Scaling Properties

  • Shilan Su,
  • Hongzhong Zhang,
  • Ziye Wang

摘要

This study investigates the scaling properties and dynamics of hashtags in tweets generated by Large Language Models (LLMs) to evaluate their performance. The study reveals that hashtags exhibit two scaling properties: Zipf’s law and Taylor’s law. However, these scaling properties have distinct characteristics compared to those observed in tweets generated by humans and natural language. Note that the relationship between the number of tokens and types does not follow Heaps’ law in contrast to human-generated tweets. The study calculates the generalized Jensen-Shannon divergence between hashtag distributions obtained at different time intervals to capture the dynamics quantitatively. The analysis shows that the decay of similarity between hashtag frequency distributions over time intervals follows a sublinear pattern similar to the dynamic behavior of hashtags in tweets generated by humans but differs from the dynamic characteristics of natural language. These findings suggest that although LLMs use hashtags differently from humans, they exhibit certain similarities in statistical properties and dynamics. The study is significant in understanding LLMs’ behavior on social media platforms and can help inform future AI text-generation strategies.