The performance of Large Language Models (LLMs) in a specific task is critically dependent upon the instruction (prompt) received. Prompt optimization aims at selecting a sequence of tokens which concatenated with the text query yields an instruction which optimizes some performance measure. Prompt optimization can be considered as a combinatorial optimization problem. The size of the combinatorial search space requires a very efficient search strategy. We propose a Bayesian Optimization method which is performed over a continues relaxation of the combinatorial search space. We consider the LLM as a black-box. Albeit preliminary and based on “vanilla” Bayesian Optimization algorithms, our experiments on several benchmark datasets show a good performance when compared against other state-of-the-art methods for prompt optimization. The numerical experiments have been conducted on three classification tasks: sentiment analysis, sentence similarity and word in context. The main performance measure is execution accuracy defined as misclassification error (0,1 loss function).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Bayesian Approach for Prompt Optimization in LLMs

  • Antonio Sabbatella,
  • Andrea Ponti,
  • Ilaria Giordani,
  • Francesco Archetti

摘要

The performance of Large Language Models (LLMs) in a specific task is critically dependent upon the instruction (prompt) received. Prompt optimization aims at selecting a sequence of tokens which concatenated with the text query yields an instruction which optimizes some performance measure. Prompt optimization can be considered as a combinatorial optimization problem. The size of the combinatorial search space requires a very efficient search strategy. We propose a Bayesian Optimization method which is performed over a continues relaxation of the combinatorial search space. We consider the LLM as a black-box. Albeit preliminary and based on “vanilla” Bayesian Optimization algorithms, our experiments on several benchmark datasets show a good performance when compared against other state-of-the-art methods for prompt optimization. The numerical experiments have been conducted on three classification tasks: sentiment analysis, sentence similarity and word in context. The main performance measure is execution accuracy defined as misclassification error (0,1 loss function).