Unlocking Imagination: Story Generation from Autistic Childrens’ Artwork Using a Fusion of Computer Vision and NLP
摘要
This research explores the potential of Natural Language Processing (NLP) techniques and computer vision to help develop the imaginative skills of autistic children by creating stories based on their artwork. This study makes use of You Only Look Once (YOLOv5) model which is a Convolutional Neural Network (CNN) to analyze and detect objects within the drawings made by autistic children, followed by NLP models such as Long Short-Term Memory (LSTM), Text-To-Text Transfer Transformer (t5), and Large Language Model (LLM) to generate narratives using detected elements. This study compares all three NLP models used to generate stories, and the best story that is generated is provided as output. The fusion of computer vision and NLP enables the automated analysis of complex details in the artwork and facilitates the generation of coherent and engaging stories. The goal is to provide a supportive tool for the caretakers, therapists, and parents that encourages cognitive development and nurtures the creative thinking of autistic kids in a personalized manner.