In recent years, with the development of Deep Neural Network (DNN), Scene Text (ST) related research has made great progress. At the same time, many methods of Scene Text Manipulation (STM) are derived. However, from the perspective of the development of deepfake, this image manipulation technology has surpassed the human eye to distinguish between real and composite images. To prevent people from the malicious application, many anti-deepfake methods have been derived. Based on the lessons learned from deepfake, since the STM has now reached a level that is indistinguishable to the naked eye, it also faces the possibility of being used by people to maliciously guide behaviors (like interfering with automatic driving to recognize road signs, forging signatures, and modifying words to cause misunderstanding semantics, etc.). As far as we know, since STM has not formed a complete manipulated dataset, we are the first to explore the identification of image manipulation. We use mainstream manipulation methods to construct a new dataset, and we designed a model specially used to identify whether the scene text image is fake. Extensive experimental comparisons with mainstream anti-deepfake and image classification methods demonstrate the effectiveness of our method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-Stream Based Scene Text Manipulation Detection Method

  • Jiefu Chen,
  • Guofeng Yi

摘要

In recent years, with the development of Deep Neural Network (DNN), Scene Text (ST) related research has made great progress. At the same time, many methods of Scene Text Manipulation (STM) are derived. However, from the perspective of the development of deepfake, this image manipulation technology has surpassed the human eye to distinguish between real and composite images. To prevent people from the malicious application, many anti-deepfake methods have been derived. Based on the lessons learned from deepfake, since the STM has now reached a level that is indistinguishable to the naked eye, it also faces the possibility of being used by people to maliciously guide behaviors (like interfering with automatic driving to recognize road signs, forging signatures, and modifying words to cause misunderstanding semantics, etc.). As far as we know, since STM has not formed a complete manipulated dataset, we are the first to explore the identification of image manipulation. We use mainstream manipulation methods to construct a new dataset, and we designed a model specially used to identify whether the scene text image is fake. Extensive experimental comparisons with mainstream anti-deepfake and image classification methods demonstrate the effectiveness of our method.