Tabular data is widely used as a common form of relational data in many industries where machine learning is applied, especially in finance, healthcare, and industry. However, traditional methods often do not fully consider the information embedded in the samples in the tables, including the meaning of the features themselves (column descriptions) as well as the contextual information. In this paper, we propose the BertTab model, which transforms a table sample into a sentence describing that sample by using an utterance template into which the category features of the table sample are populated. The sentences describing the sample are then converted into powerful contextual embeddings using the pre-trained Bert model. Finally, the context-embedded features are fused with the original features to obtain richer and more complete features of the sample, thus achieving higher performance. We evaluated the model on three datasets. Compared to the benchmark model, BertTab improves the AUC-ROC, AUC-PR, and accuracy by an average of 2.10, 4.43, and 0.48% on the three datasets, respectively. The ablation experiments demonstrate the positive effect of introducing category feature column descriptions and considering sample category feature contexts with the fusion of raw features on model effectiveness improvement.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BertTab: Table Learning with Feature Descriptions and Context

  • Meng Xie,
  • Haibing An,
  • Songjun Han,
  • Jiangtao Mao,
  • Ying Jiang,
  • Jiajia Wang

摘要

Tabular data is widely used as a common form of relational data in many industries where machine learning is applied, especially in finance, healthcare, and industry. However, traditional methods often do not fully consider the information embedded in the samples in the tables, including the meaning of the features themselves (column descriptions) as well as the contextual information. In this paper, we propose the BertTab model, which transforms a table sample into a sentence describing that sample by using an utterance template into which the category features of the table sample are populated. The sentences describing the sample are then converted into powerful contextual embeddings using the pre-trained Bert model. Finally, the context-embedded features are fused with the original features to obtain richer and more complete features of the sample, thus achieving higher performance. We evaluated the model on three datasets. Compared to the benchmark model, BertTab improves the AUC-ROC, AUC-PR, and accuracy by an average of 2.10, 4.43, and 0.48% on the three datasets, respectively. The ablation experiments demonstrate the positive effect of introducing category feature column descriptions and considering sample category feature contexts with the fusion of raw features on model effectiveness improvement.