<p>Libraries are custodians of a written heritage that is both voluminous and fragile, and needs to be regularly monitored due to the sensitivity of book materials. For most libraries, it is humanly impossible to individually monitor each book. To address this challenge, we explored the potential of deep learning to automatically gather precise data on books exhibiting dangerous structural alterations on their spines. Our research focused on the “Parlement de Paris” archives and two databases of photographs of bindings were captured. Annotations were performed distinguishing bindings from the background, labeling text, and creating binary masks corresponding to different types of alterations. The detection pipeline involved instance segmentation using Mask R-CNN to separate bookshelf images into individual binding photographs. Alteration classification was performed using a Vision Transformer (ViT) pre-trained on ImageNet, with the model fine-tuned for multilabel classification. While the detection pipeline demonstrated high performance in label detection and text recognition, challenges included underestimation of binding surface area and recognition of some alterations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep learning-based pipeline for the conservation assessment of bindings in archives and libraries

  • Valérie Lee-Gouet,
  • Lahcen Yamoun,
  • Zacharie Rodière,
  • Camille Simon Chane,
  • Michel Jordan,
  • Julien Longhi,
  • David Picard

摘要

Libraries are custodians of a written heritage that is both voluminous and fragile, and needs to be regularly monitored due to the sensitivity of book materials. For most libraries, it is humanly impossible to individually monitor each book. To address this challenge, we explored the potential of deep learning to automatically gather precise data on books exhibiting dangerous structural alterations on their spines. Our research focused on the “Parlement de Paris” archives and two databases of photographs of bindings were captured. Annotations were performed distinguishing bindings from the background, labeling text, and creating binary masks corresponding to different types of alterations. The detection pipeline involved instance segmentation using Mask R-CNN to separate bookshelf images into individual binding photographs. Alteration classification was performed using a Vision Transformer (ViT) pre-trained on ImageNet, with the model fine-tuned for multilabel classification. While the detection pipeline demonstrated high performance in label detection and text recognition, challenges included underestimation of binding surface area and recognition of some alterations.