Solving Rate of Mathematical Olympiad Word Problems by Artificial Intelligences
摘要
Large Language Models (LLMs) have made significant strides in natural language processing, but their effectiveness in solving mathematical word problems remains underexplored. This study examines the performance of various AI models, including GPT-4o and OpenAI’s o1-preview, in solving word problems from the Slovak Mathematical Olympiad. While OpenAI reported an 86% success rate for o1-preview on International Mathematical Olympiad tasks, our results, focused on Slovak word problems, show lower rates: 43% for o1-preview and 14% for GPT-4o. Despite these discrepancies, ChatGPT o1-preview significantly outperforms previous models. The study highlights the models’ current limitations and suggests that with further development, AI could enhance mathematical problem-solving.