An open benchmark for oracle bone rubbing image retrieval
摘要
Oracle bone inscriptions provide critical insights into ancient Chinese history. However, the retrieval and analysis of inscription rubbings remain challenging due to fragmentation, weathering, and non-standardized character forms. These challenges fundamentally limit the applicability of conventional image retrieval methods, an issue exacerbated by the lack of large-scale annotated datasets. To tackle these challenges, we introduce the first dataset and a Multi-step Strategy for Homologous Rubbing Retrieval (MSHRR). MSHRR employs a three-stage pipeline integrating character extraction, cross-rubbing matching, and similarity scoring, bypassing Optical Character Recognition (OCR) dependencies. This novel framework outperforms state-of-the-art methods in handling glyph structures through its morphology-aware paradigm. More importantly, MSHRR has found 276 new homologous sets, accounting for over 10% of documented cases in twenty years. Our benchmark also offers a reproducible evaluation framework for computational archeology and reveals new historical connections.