<p>Large data often has characteristics such as large volume, high dimensionality, multi-temporal, noise, uncertainty or approximation, etc. Large data is also often difficult to process and store centrally due to high requirements for equipment infrastructure. These factors can significantly affect the data analysis process. This paper proposes a collaborative learning model based on an interval type-2 fuzzy set (IT2FS) and the semi-supervised clustering method for large data analysis problems. Our method aims to solve the problem of large data with limited labeled data. The collaborative learning model not only allows decentralized data processing but also allows the sharing of data analysis results across data sites, which can help improve the accuracy of data analysis results. With the data representation based on IT2FSs, it is possible to reduce the influence of the uncertainty within the dataset. By combining the semi-supervised technique with the limited amount of labeled data to guide the clustering, this approach can help improve the accuracy of the clustering results. Experiments on large real datasets downloaded from the UCI library and remote sensing image data demonstrate that the proposed method yields significantly better clustering results than several previously proposed methods. This result also demonstrates the potential of collaborative clustering models in analyzing large and distributed datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A collaborative learning model using the semi-supervised method and the interval type-2 fuzzy set for large data analysis

  • Viet Duc Do,
  • Dinh Sinh Mai,
  • Long Thanh Ngo

摘要

Large data often has characteristics such as large volume, high dimensionality, multi-temporal, noise, uncertainty or approximation, etc. Large data is also often difficult to process and store centrally due to high requirements for equipment infrastructure. These factors can significantly affect the data analysis process. This paper proposes a collaborative learning model based on an interval type-2 fuzzy set (IT2FS) and the semi-supervised clustering method for large data analysis problems. Our method aims to solve the problem of large data with limited labeled data. The collaborative learning model not only allows decentralized data processing but also allows the sharing of data analysis results across data sites, which can help improve the accuracy of data analysis results. With the data representation based on IT2FSs, it is possible to reduce the influence of the uncertainty within the dataset. By combining the semi-supervised technique with the limited amount of labeled data to guide the clustering, this approach can help improve the accuracy of the clustering results. Experiments on large real datasets downloaded from the UCI library and remote sensing image data demonstrate that the proposed method yields significantly better clustering results than several previously proposed methods. This result also demonstrates the potential of collaborative clustering models in analyzing large and distributed datasets.