An implicit layout-aware transformer for full-page end-to-end optical music recognition
摘要
The digital preservation and accessibility of musical manuscripts present a significant challenge in musical heritage conservation. Optical Music Recognition (OMR) is a key solution for automating the transcription of sheet music into machine-readable formats, typically relying on multi-stage processing due to the complexity of musical notation. Although the trend in OMR is to reduce the number of processing steps, the latest single-step approaches lose layout information from the original image, which is crucial for accurately interpreting musical documents in practical applications. This paper presents the Layout-Aware Sheet Music Transformer, a novel method based on a single-step transcription strategy that leverages transformer attention mechanisms to extract layout information without explicit annotations. By dynamically analyzing internal model representations during direct image-to-sequence transcription, our approach extracts nuanced structural insights without requiring manual annotations, eliminating manual layout annotations while offering deeper insights into document structures that more faithfully reflect the underlying musical meaning. Experiments in four well-known OMR datasets show that our proposal is both superior in transcription and competitive in layout extraction, making it a solid candidate for a new state of the art in the field.