Large Language Models (LLMs) have made significant strides in natural language processing, but their effectiveness in solving mathematical word problems remains underexplored. This study examines the performance of various AI models, including GPT-4o and OpenAI’s o1-preview, in solving word problems from the Slovak Mathematical Olympiad. While OpenAI reported an 86% success rate for o1-preview on International Mathematical Olympiad tasks, our results, focused on Slovak word problems, show lower rates: 43% for o1-preview and 14% for GPT-4o. Despite these discrepancies, ChatGPT o1-preview significantly outperforms previous models. The study highlights the models’ current limitations and suggests that with further development, AI could enhance mathematical problem-solving.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Solving Rate of Mathematical Olympiad Word Problems by Artificial Intelligences

  • Jana Fialová,
  • Roman Horváth

摘要

Large Language Models (LLMs) have made significant strides in natural language processing, but their effectiveness in solving mathematical word problems remains underexplored. This study examines the performance of various AI models, including GPT-4o and OpenAI’s o1-preview, in solving word problems from the Slovak Mathematical Olympiad. While OpenAI reported an 86% success rate for o1-preview on International Mathematical Olympiad tasks, our results, focused on Slovak word problems, show lower rates: 43% for o1-preview and 14% for GPT-4o. Despite these discrepancies, ChatGPT o1-preview significantly outperforms previous models. The study highlights the models’ current limitations and suggests that with further development, AI could enhance mathematical problem-solving.