This paper presents two experiments testing the capabilities of the models GPT-4-vision-preview and GPT-4o, to transcribe images of randomly generated number-strings. We noticed a tendency for longer number strings to be read inaccurately while transcribing invoices using the models, so we created an experiment with images with minimal noise to see what amount of digits these models are able to transcribe with full accuracy. We further tested weather the mistakes that occur in transcribing longer strings are reoccurring when the images are re-tested. We analyzed the character of the mistakes made, and weather any digits are over represented among the ones involved in mistakes. We found, among more results described in the paper, that the models are 100% accurate up to 75 digits per image, that the same mistakes are reoccurring when the same image is rerun, and that hallucinations are only 23% of all mistakes made by the models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hallucinations and Training-Data Bias: Results from Two Number Transcription Experiments Using GPT Models

  • Nemi Pelgrom,
  • Johan Hangelbäck,
  • Morgan Ericsson,
  • Jonas Nordqvist,
  • Håkan Grahn

摘要

This paper presents two experiments testing the capabilities of the models GPT-4-vision-preview and GPT-4o, to transcribe images of randomly generated number-strings. We noticed a tendency for longer number strings to be read inaccurately while transcribing invoices using the models, so we created an experiment with images with minimal noise to see what amount of digits these models are able to transcribe with full accuracy. We further tested weather the mistakes that occur in transcribing longer strings are reoccurring when the images are re-tested. We analyzed the character of the mistakes made, and weather any digits are over represented among the ones involved in mistakes. We found, among more results described in the paper, that the models are 100% accurate up to 75 digits per image, that the same mistakes are reoccurring when the same image is rerun, and that hallucinations are only 23% of all mistakes made by the models.