Large Language Models (LLMs) are increasingly used to solve machine learning tasks on tabular data, such as classification tasks. Using standard benchmarks, recent studies have shown impressive performance for such tasks. However, in this paper we study a major concern with the use of standard benchmarks: LLMs may have also been trained using these benchmarks, that is, these LLMs are contaminated. While previous work mostly focused on the GPT models, we propose a methodology to evaluate whether a large range of LLMs is contaminated. We propose new tests to detect contamination with tabular data over two aspects: knowledge and memorization. We also design an algorithm to parse the answers of the LLMs and detect hints of contamination. Our experiments conclude that the bigger and more closed-source an LLM is, the more likely it is contaminated.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detection of Large Language Model Contamination with Tabular Data

  • Benoît Ronval,
  • Pierre Dupont,
  • Siegfried Nijssen

摘要

Large Language Models (LLMs) are increasingly used to solve machine learning tasks on tabular data, such as classification tasks. Using standard benchmarks, recent studies have shown impressive performance for such tasks. However, in this paper we study a major concern with the use of standard benchmarks: LLMs may have also been trained using these benchmarks, that is, these LLMs are contaminated. While previous work mostly focused on the GPT models, we propose a methodology to evaluate whether a large range of LLMs is contaminated. We propose new tests to detect contamination with tabular data over two aspects: knowledge and memorization. We also design an algorithm to parse the answers of the LLMs and detect hints of contamination. Our experiments conclude that the bigger and more closed-source an LLM is, the more likely it is contaminated.