Dark patterns, i.e., deceptive design patterns that employ manipulative strategies to deceive or force online users to take a decision against their interests, are widely available in online world. However, the variety and quantity of structured and labelled dark pattern datasets, which are critical for automated dark pattern detection, particularly for AI-based detection models, are limited. In this study, we leverage Large Language Models’ (LLM) sophisticated text data generation ability and propose a dark pattern text data augmentation method by utilizing a state of art open source language model and multi-agents framework, which has generator and controller models. Evaluation of the augmentation demonstrates that while increasing the data size, our proposal-based augmented data preserves the same dark pattern characteristics of the source data and maintains its diversity. We set forth that dark pattern text data can be generated even based on a few examples via prompt engineering techniques on the LLMs. We also show that our augmented data can be used to fine-tune pre-trained language models using Low-Rank Adaptation to enhance their robustness in detecting dark patterns.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Augmenting Dark Patterns Text Data by Leveraging Large Language Models: A Multi-agent Framework and Parameter-Efficient Fine-Tuning

  • Emre Kocyigit,
  • Davide Liga,
  • Gabriele Lenzini

摘要

Dark patterns, i.e., deceptive design patterns that employ manipulative strategies to deceive or force online users to take a decision against their interests, are widely available in online world. However, the variety and quantity of structured and labelled dark pattern datasets, which are critical for automated dark pattern detection, particularly for AI-based detection models, are limited. In this study, we leverage Large Language Models’ (LLM) sophisticated text data generation ability and propose a dark pattern text data augmentation method by utilizing a state of art open source language model and multi-agents framework, which has generator and controller models. Evaluation of the augmentation demonstrates that while increasing the data size, our proposal-based augmented data preserves the same dark pattern characteristics of the source data and maintains its diversity. We set forth that dark pattern text data can be generated even based on a few examples via prompt engineering techniques on the LLMs. We also show that our augmented data can be used to fine-tune pre-trained language models using Low-Rank Adaptation to enhance their robustness in detecting dark patterns.