Performance Evaluation of ChatGPT on BITSAT: Engineering Entrance Examination
摘要
ChatGPT, a powerful natural language processing tool, is an Artificial Intelligence (AI) language model designed to communicate with humans. The education field has achieved a new perspective with AI playing a very important role. AI has opened avenues to comprehend subjects, get answers to questions, and receive recommendations and problem-solving solutions. The present study aims to assess the performance outcome of ChatGPT in solving Birla Institute of Technology and Science Admission Test (BITSAT). BITSAT is an engineering entrance examination conducted in India. An adjudication criterion based on Accuracy, Concordance, and Insight is implemented for performance evaluation of ChatGPT responses. A dataset of questions and answers for subjects including physics, chemistry, English proficiency, logical reasoning, and mathematics is generated from the previous two years’ exams. ChatGPT responses pertaining to English have the highest accuracy (72%), concordance (52%), and insight (68%), while least with physics at 30%, 28%, and 30%, respectively. The variation in the outcomes can be attributed to factors such as lack of domain expertise, and limited ability to understand the context and generate original insights. It is therefore imperative for academia to be abreast of the strengths and weaknesses of ChatGPT, so as to make an effective utilization of ChatGPT as a tool in the field of academics and research.