<p>Puzzle-solving is a problem having applications for instance in archaeology and cultural heritage. Proposed solutions often suffer from a performance loss when the pieces are eroded, a characteristic that is pervasive across various use—cases such as frescoes reconstruction. Most approaches divide the problem into two fundamental phases: discriminating and then positioning the pieces. We focus on the case of puzzles with square pieces, without any missing or extraneous pieces, and we introduce the first two-step deep learning solution capable of efficiently solving puzzles, from discrimination to piece placement, while remaining robust to erosion. In the context of permutation learning, we propose to use transformers to determine the correct placement of the pieces and an encoder that uses the information at the edge of the pieces. This method sets a new state of the art, achieving a significant performance gain, and introduces a new approach for learning similarity functions in the context of puzzle solving.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Solving jigsaw puzzles with vision transformers

  • Gaël Heck,
  • Nicolas Lermé,
  • Sylvie Le Hégarat-Mascle

摘要

Puzzle-solving is a problem having applications for instance in archaeology and cultural heritage. Proposed solutions often suffer from a performance loss when the pieces are eroded, a characteristic that is pervasive across various use—cases such as frescoes reconstruction. Most approaches divide the problem into two fundamental phases: discriminating and then positioning the pieces. We focus on the case of puzzles with square pieces, without any missing or extraneous pieces, and we introduce the first two-step deep learning solution capable of efficiently solving puzzles, from discrimination to piece placement, while remaining robust to erosion. In the context of permutation learning, we propose to use transformers to determine the correct placement of the pieces and an encoder that uses the information at the edge of the pieces. This method sets a new state of the art, achieving a significant performance gain, and introduces a new approach for learning similarity functions in the context of puzzle solving.