<p>AI technology’s rapid advancements have led to its widespread adoption in diverse industries, such as support agents, healthcare, education, software development and finance because such AI technologies has the ability to generated human-like text to solve a complex user’s problem. Nevertheless, the effectiveness and success of AI models greatly depends on large language model that contain extensive, varied, and top-notch datasets and social network platforms for both training and evaluation purposes that is sourced from web and directly form user interaction that will also contain sensitive information. However, this power capabilities and progress have raised significant concerns related data security and privacy, particularly regarding protection of user data including social network during training and deployment phase. Large language models process and generate huge amount of data which have the potential to unintentionally memorise and repeat personal sensitive information like location data, social security number, and phone numbers which are shared by users that can lead to pose a risk of data leaking and degrade the accuracy of model training and information. This paper will perform a comprehensive and systematic review to examine the fundamentals principle of large language model with particularly concentration on data privacy concern associated with large language model during training data and potential biases. Furthermore, we analysis and evaluated the current level of vulnerabilities, privacy challenges, and examine new security and privacy attacks targeting LLMs that cause of misuse and information leakage. Additionally, we proposed a set of best practices aimed to protect user data. Recognising the significance of conducting such research on protecting privacy and improving security is crucial because it ensures the continuous progress and public trust in LLM technology.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Securing social network user data in large language model deployments: challenges and best practices

  • Nasir Ahmad Jalali,
  • Chen Hongsong

摘要

AI technology’s rapid advancements have led to its widespread adoption in diverse industries, such as support agents, healthcare, education, software development and finance because such AI technologies has the ability to generated human-like text to solve a complex user’s problem. Nevertheless, the effectiveness and success of AI models greatly depends on large language model that contain extensive, varied, and top-notch datasets and social network platforms for both training and evaluation purposes that is sourced from web and directly form user interaction that will also contain sensitive information. However, this power capabilities and progress have raised significant concerns related data security and privacy, particularly regarding protection of user data including social network during training and deployment phase. Large language models process and generate huge amount of data which have the potential to unintentionally memorise and repeat personal sensitive information like location data, social security number, and phone numbers which are shared by users that can lead to pose a risk of data leaking and degrade the accuracy of model training and information. This paper will perform a comprehensive and systematic review to examine the fundamentals principle of large language model with particularly concentration on data privacy concern associated with large language model during training data and potential biases. Furthermore, we analysis and evaluated the current level of vulnerabilities, privacy challenges, and examine new security and privacy attacks targeting LLMs that cause of misuse and information leakage. Additionally, we proposed a set of best practices aimed to protect user data. Recognising the significance of conducting such research on protecting privacy and improving security is crucial because it ensures the continuous progress and public trust in LLM technology.