<p>The storage, processing, and sharing of historical documents are an essential part of cultural heritage research that needs efficient storage, smart processing, and teamwork. The research primarily proposes an edge-based historical document analysis framework using a TCO-optimized Long Short-Term Memory (TCO-Vanilla LSTM) model for efficient recognition. Binary classification was employed as an initial filtering mechanism to distinguish historical documents from non-relevant records, supporting streamlined large-scale processing. A high-resolution dataset consisting of 40 digitized historical document images was used, sourced from libraries and archival repositories. The system was implemented using Python. On edge nodes, preprocessing methods, such as normalization, artifact elimination, noise elimination, and bilateral filtering, are used to improve the quality of images, the edge information, and the textual information, and the transmission overhead. The system uses a distributed storage architecture that divides and copies documents across edge nodes, geographically distributed, fault-tolerant, scalable, and has low latency. The TCO optimizes hyperparameters of the model, enhancing performance, stability, and reducing latency during training and inference. The Vanilla Long Short-Term Memory (LSTM) network is used to process sequential data from historical documents, enabling accurate classification and recognition of text patterns. The experimental findings show that the offered TCO-Vanilla LSTM model has a better delay performance, accuracy of 0.9416, and an F1-score of 0.92, compared to the traditional cloud-based models in terms of scalability, efficiency, and collaborative document analysis.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distributed storage and collaborative computing platform for historical documents under edge intelligence architecture

  • Yue Wang

摘要

The storage, processing, and sharing of historical documents are an essential part of cultural heritage research that needs efficient storage, smart processing, and teamwork. The research primarily proposes an edge-based historical document analysis framework using a TCO-optimized Long Short-Term Memory (TCO-Vanilla LSTM) model for efficient recognition. Binary classification was employed as an initial filtering mechanism to distinguish historical documents from non-relevant records, supporting streamlined large-scale processing. A high-resolution dataset consisting of 40 digitized historical document images was used, sourced from libraries and archival repositories. The system was implemented using Python. On edge nodes, preprocessing methods, such as normalization, artifact elimination, noise elimination, and bilateral filtering, are used to improve the quality of images, the edge information, and the textual information, and the transmission overhead. The system uses a distributed storage architecture that divides and copies documents across edge nodes, geographically distributed, fault-tolerant, scalable, and has low latency. The TCO optimizes hyperparameters of the model, enhancing performance, stability, and reducing latency during training and inference. The Vanilla Long Short-Term Memory (LSTM) network is used to process sequential data from historical documents, enabling accurate classification and recognition of text patterns. The experimental findings show that the offered TCO-Vanilla LSTM model has a better delay performance, accuracy of 0.9416, and an F1-score of 0.92, compared to the traditional cloud-based models in terms of scalability, efficiency, and collaborative document analysis.