<p>In the modern world most of the users are using Google Gemini AI. To use Google’s Gemini AI efficiently we need the best prompting way to get the information. Google’s Gemini AI is well known for its conversational nature, but before applying it to high stake sectors like HealthCare, its responses need thorough and systematic evaluation. Our research targets the vital gap of not having a solid quantitative and qualitative framework for evaluating responses of Gemini based on different prompt structures. Our research mainly takes three types of prompts into consideration: long, short and structured in the healthcare domain. For effective evaluation, a hybrid method was taken into consideration combining qualitative factors like relevance, accuracy, depth, creativity and clarity with quantitative metrics like sufficiency of word count, keyword coverage, reliability and factuality score. Our results prove that Long prompts generally yield a better set of results compared to other forms of prompts closely followed by Structured prompt which has a major drawback of excessive verbosity. We concluded the work by exploring the best prompt structure to yield relevant medically accurate results from Gemini while also acknowledging limitations such as model non-determinism, continuous model updates, and variability in response consistency across iterations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Strategic framework design for prompt engineering analysis: GEMINI

  • Ramkumar Sivasakthivel,
  • Rupsha Das,
  • Swapnil Roy,
  • Gobinath Ramar,
  • N. Rajasekaran,
  • R. Stephen

摘要

In the modern world most of the users are using Google Gemini AI. To use Google’s Gemini AI efficiently we need the best prompting way to get the information. Google’s Gemini AI is well known for its conversational nature, but before applying it to high stake sectors like HealthCare, its responses need thorough and systematic evaluation. Our research targets the vital gap of not having a solid quantitative and qualitative framework for evaluating responses of Gemini based on different prompt structures. Our research mainly takes three types of prompts into consideration: long, short and structured in the healthcare domain. For effective evaluation, a hybrid method was taken into consideration combining qualitative factors like relevance, accuracy, depth, creativity and clarity with quantitative metrics like sufficiency of word count, keyword coverage, reliability and factuality score. Our results prove that Long prompts generally yield a better set of results compared to other forms of prompts closely followed by Structured prompt which has a major drawback of excessive verbosity. We concluded the work by exploring the best prompt structure to yield relevant medically accurate results from Gemini while also acknowledging limitations such as model non-determinism, continuous model updates, and variability in response consistency across iterations.