This paper proposes a semantically aware model for document annotation. The model encompasses incremental informative term discovery followed by generation of knowledge graphs through Google Knowledge Graph API, and staged incremental auxiliary knowledge addition using the Wikidata standard knowledge store repository. Further, ontologies are generated from the knowledge instances using the OntoCollab tool. The proposed framework then employs an encoder–decoder transformer model to classify the highly voluminous metadata which is generated from an RDF distiller applied onto the generated ontologies. The encompassment of a biologically inspired clonal selection algorithm for metaheuristic optimization under adaptive pointwise mutual information measure as a criteria function helps yield the best-in-class facets for annotation of agro-based documents. On the other end of the pipeline pertaining to horticultural documents, quantitative semantics-oriented computation of SimRank and CoSimRank on the entities yielded by OntoCollab, Wikidata, Google Knowledge Graph API, and Term Frequency–Inverse Document Frequency are used with empirically decided thresholds and step deviance measures to rank and finalize the most relevant facets for annotation of horticultural documents. The proposed framework achieves much more efficacious performance thereby making it the best-in-class model for document annotation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SDPSH: Semantically Aware Annotation of Documents Pertaining to Sustainable Horticulture and Agro-Based Farming

  • Archit Chadalawada,
  • Gerard Deepak

摘要

This paper proposes a semantically aware model for document annotation. The model encompasses incremental informative term discovery followed by generation of knowledge graphs through Google Knowledge Graph API, and staged incremental auxiliary knowledge addition using the Wikidata standard knowledge store repository. Further, ontologies are generated from the knowledge instances using the OntoCollab tool. The proposed framework then employs an encoder–decoder transformer model to classify the highly voluminous metadata which is generated from an RDF distiller applied onto the generated ontologies. The encompassment of a biologically inspired clonal selection algorithm for metaheuristic optimization under adaptive pointwise mutual information measure as a criteria function helps yield the best-in-class facets for annotation of agro-based documents. On the other end of the pipeline pertaining to horticultural documents, quantitative semantics-oriented computation of SimRank and CoSimRank on the entities yielded by OntoCollab, Wikidata, Google Knowledge Graph API, and Term Frequency–Inverse Document Frequency are used with empirically decided thresholds and step deviance measures to rank and finalize the most relevant facets for annotation of horticultural documents. The proposed framework achieves much more efficacious performance thereby making it the best-in-class model for document annotation.