Nowadays, vast images are generated daily as we capture, transfer, and receive them in various sectors. Images have become a pivotal component of data in many different industries, contributing to decision-making, documentation, and artistic expression. However, verifying image authenticity has become more complex with the widespread availability of sophisticated software and tools that enable image alteration. As a result, determining whether an image is original or manipulated has become a complex task. In this paper, we propose an enhanced Transformer architecture to classify between original and manipulated images by using their metadata and EXIF data. Two datasets are built to train the framework. Each dataset carries metadata and EXIF data of original and manipulated images, respectively. An augmentation technique has been applied to ensure dataset balance and robustness. The proposed framework uses a parallel multi-head attention mechanism, which speeds up convergence throughout the training process and results in more efficient model learning. This versatile proposed framework can perform on different image formats such as JPG/JPEG, PNG, and BMP, highlighting its adaptability and real-world applicability. This framework has achieved 96.42% accuracy, showing its potentiality and capability to distinguish between original and manipulated images in this digital age.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Image Forensics with Transformer: A Multi-head Attention Approach for Robust Metadata Analysis

  • Md. Appel Mahmud Pranto,
  • Nafiz Al Asad,
  • Mohammad Abu Yousuf,
  • Mohammed Nasir Uddin,
  • Mohammad Ali Moni

摘要

Nowadays, vast images are generated daily as we capture, transfer, and receive them in various sectors. Images have become a pivotal component of data in many different industries, contributing to decision-making, documentation, and artistic expression. However, verifying image authenticity has become more complex with the widespread availability of sophisticated software and tools that enable image alteration. As a result, determining whether an image is original or manipulated has become a complex task. In this paper, we propose an enhanced Transformer architecture to classify between original and manipulated images by using their metadata and EXIF data. Two datasets are built to train the framework. Each dataset carries metadata and EXIF data of original and manipulated images, respectively. An augmentation technique has been applied to ensure dataset balance and robustness. The proposed framework uses a parallel multi-head attention mechanism, which speeds up convergence throughout the training process and results in more efficient model learning. This versatile proposed framework can perform on different image formats such as JPG/JPEG, PNG, and BMP, highlighting its adaptability and real-world applicability. This framework has achieved 96.42% accuracy, showing its potentiality and capability to distinguish between original and manipulated images in this digital age.