Hallucinations and Training-Data Bias: Results from Two Number Transcription Experiments Using GPT Models
摘要
This paper presents two experiments testing the capabilities of the models GPT-4-vision-preview and GPT-4o, to transcribe images of randomly generated number-strings. We noticed a tendency for longer number strings to be read inaccurately while transcribing invoices using the models, so we created an experiment with images with minimal noise to see what amount of digits these models are able to transcribe with full accuracy. We further tested weather the mistakes that occur in transcribing longer strings are reoccurring when the images are re-tested. We analyzed the character of the mistakes made, and weather any digits are over represented among the ones involved in mistakes. We found, among more results described in the paper, that the models are 100% accurate up to 75 digits per image, that the same mistakes are reoccurring when the same image is rerun, and that hallucinations are only 23% of all mistakes made by the models.