The basic goal of this work is to improve record linkage in uncertain datasets by utilizing distance measurement and probabilistic functions as its two primary methods. This will be accomplished by improving record linkage in uncertain datasets. In most cases, a semantic vector, in which each component has its own weighted value, is utilized to perform the calculation that determines an estimate of the distance that exists between two points. The best agreement principle and the distance function are applied together to accomplish the task of deciding which ontological routes can be pursued. This is done so that we can determine which options are available. As a result of this, the incongruent pairs are excluded from the calculation of distance, which results in an improvement in the precision of record linking. Conventional approaches such as supervised learning, probabilistic models, and graph linking are used to evaluate the suggested strategy in relation to the present state of the art. This is done to determine how well the proposed strategy stacks up. The datasets NCBI GenBank, Gene Ontology, and the SwissProt dataset are utilized for evaluations of precision, recall, and F-measure, in addition to the examination of additional metrics. According to the findings, the suggested method outperforms the present state of the art in terms of precision, recall rate, and F-measure, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Record Linkage as a Pre-processing Tool in Big Data Applications Using Machine Learning

  • Kayalvizhi Subramanian,
  • Gunasekar Thangarasu,
  • Nattar Kannan Kalliappan

摘要

The basic goal of this work is to improve record linkage in uncertain datasets by utilizing distance measurement and probabilistic functions as its two primary methods. This will be accomplished by improving record linkage in uncertain datasets. In most cases, a semantic vector, in which each component has its own weighted value, is utilized to perform the calculation that determines an estimate of the distance that exists between two points. The best agreement principle and the distance function are applied together to accomplish the task of deciding which ontological routes can be pursued. This is done so that we can determine which options are available. As a result of this, the incongruent pairs are excluded from the calculation of distance, which results in an improvement in the precision of record linking. Conventional approaches such as supervised learning, probabilistic models, and graph linking are used to evaluate the suggested strategy in relation to the present state of the art. This is done to determine how well the proposed strategy stacks up. The datasets NCBI GenBank, Gene Ontology, and the SwissProt dataset are utilized for evaluations of precision, recall, and F-measure, in addition to the examination of additional metrics. According to the findings, the suggested method outperforms the present state of the art in terms of precision, recall rate, and F-measure, respectively.