A Plug-and-Play and Invisible Multi-Bit Watermarking Scheme for Deep Neural Networks
摘要
With the widespread deployment of deep neural networks in business and industry, protecting model intellectual property has become increasingly critical. Traditional black-box watermarking methods often neglect trigger-set invisibility, leaving them vulnerable to detection and removal. Although some recent methods improve invisibility, they typically require complex fine-tuning or retraining, increasing costs and harming model performance. Moreover, existing multi-bit watermarking schemes incur significant storage and computational overhead due to their reliance on large trigger sets. To address these limitations, we propose a novel plug-and-play invisible multi-bit watermarking scheme that jointly trains a visual consistency model and an auxiliary model to generate invisible trigger samples. By embedding the auxiliary model into the target model, we inject watermark information directly into the logits output and final results, without fine-tuning. Furthermore, the same trigger set can yield multiple responses, eliminating the need for multiple trigger sets. Experimental results on GTSRB, CIFAR-100, and ImageNet demonstrate that our method achieves watermark embedding with negligible performance loss, accurate ownership verification, and robustness against various attacks including model fine-tuning, pruning, and model stealing.