Learning-Based Sub-image Retrieval in Historical Document Images
摘要
The goal of this paper is to propose an unsupervised learning-based framework in order to deal with any kind of one-shot object detection scenario, focusing on the tasks of sub-image retrieval and pattern spotting in historical document images. Taking in an arbitrary object/pattern query from users, the proposed framework should be able to retrieve images containing it, as well as localising each occurrence within the images. A major difficulty is the lack of any training data. Three contributions are thus presented: (1) a novel model architecture dubbed OS-DETR, capable of adapting to various tasks by simply swapping training data, (2) a completely unsupervised synthetic data generation process, easily applicable to many data-limited domains, and (3) a set of training strategies catered to boost the model’s generalisation capabilities. The result is a framework that yields a strong baseline for learning-based approaches applied to sub-image retrieval and pattern spotting.