It is a fundamental challenge to evaluate whether a model can truly capture the meaning of sentences. Evaluation of whether a model well captures the meaning of individual words, however, can be effectively achieved by analyzing whether the model encodes words in a vector space where semantically similar words form clusters. Inspired by this approach, we propose the Sentence-Space Metrics (SSM) to evaluate model interpretation of sentences, and the sentence space is constructed based on the pairwise entailment relationships between all sentence pairs within a sentence pool. We use three metrics to evaluate a sentence space, i.e., (1) sparsity, (2) clustering of related sentences, and (3) similarity with the sentence space measured from humans. The SSM is applied to evaluate 20 models, including ChatGPT, 18 BERT-family models fine-tuned for Natural Language Inference (NLI) task, as well as SimCSE, a sentence representation model. The SSM reveals dramatic differences among models: Although all models achieve high accuracy on standard NLI datasets such as MNLI, none of them mirrors the human behavior under the SSM. These results demonstrate that, compared with traditional accuracy measures, the SSM considers pairwise relationships between hundreds of sentences and therefore provide a more fine-grained evaluation of model interpretation of sentences.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sentence-Space Metrics (SSM) for the Evaluation of Sentence Comprehension

  • Jieyu Lin,
  • Honghua Chen,
  • Nai Ding

摘要

It is a fundamental challenge to evaluate whether a model can truly capture the meaning of sentences. Evaluation of whether a model well captures the meaning of individual words, however, can be effectively achieved by analyzing whether the model encodes words in a vector space where semantically similar words form clusters. Inspired by this approach, we propose the Sentence-Space Metrics (SSM) to evaluate model interpretation of sentences, and the sentence space is constructed based on the pairwise entailment relationships between all sentence pairs within a sentence pool. We use three metrics to evaluate a sentence space, i.e., (1) sparsity, (2) clustering of related sentences, and (3) similarity with the sentence space measured from humans. The SSM is applied to evaluate 20 models, including ChatGPT, 18 BERT-family models fine-tuned for Natural Language Inference (NLI) task, as well as SimCSE, a sentence representation model. The SSM reveals dramatic differences among models: Although all models achieve high accuracy on standard NLI datasets such as MNLI, none of them mirrors the human behavior under the SSM. These results demonstrate that, compared with traditional accuracy measures, the SSM considers pairwise relationships between hundreds of sentences and therefore provide a more fine-grained evaluation of model interpretation of sentences.