In semi-supervised and active learning, it is assumed that human experts can provide a set of reliably labeled instances. Crowdworking is used increasingly as a business model to acquire such labels by exploiting the wisdom of the crowd. A less investigated application area is a posteriori labeling of medical archives and of entries from population-based studies with diagnoses of diseases that were not recorded when the studies were conducted. Such labeling would make it possible for epidemiologists and clinical researchers to learn about the onset and evolution of diseases from earlier data collections. Data labeling is a complex process, especially for humans. Labels are still needed, but a completely automated approach to data acquisition demands identification of hallucinations, and this can only be done by humans. Methods such as few-shot learning have limitations, e.g. poor data at the beginning to learn the model. We investigate the potential of acquiring labels through the pairwise comparison of instances, and then propagating these labels in a semi-supervised way. For this purpose, we use structured data with hepatic steatosis as outcome and show that its possible, based on very few data to classify this data correctly.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semi-supervised Learning with Pairwise Instance Comparisons for Medical Instance Classification

  • Anne Rother,
  • Till Ittermann,
  • Myra Spiliopoulou

摘要

In semi-supervised and active learning, it is assumed that human experts can provide a set of reliably labeled instances. Crowdworking is used increasingly as a business model to acquire such labels by exploiting the wisdom of the crowd. A less investigated application area is a posteriori labeling of medical archives and of entries from population-based studies with diagnoses of diseases that were not recorded when the studies were conducted. Such labeling would make it possible for epidemiologists and clinical researchers to learn about the onset and evolution of diseases from earlier data collections. Data labeling is a complex process, especially for humans. Labels are still needed, but a completely automated approach to data acquisition demands identification of hallucinations, and this can only be done by humans. Methods such as few-shot learning have limitations, e.g. poor data at the beginning to learn the model. We investigate the potential of acquiring labels through the pairwise comparison of instances, and then propagating these labels in a semi-supervised way. For this purpose, we use structured data with hepatic steatosis as outcome and show that its possible, based on very few data to classify this data correctly.