Verbal Representation of Object Collision Prediction Based on Physical CommonSense Knowledge
摘要
In recent years, many prediction models for real-world objects have been proposed. Most research on prediction produces prediction results from visual predictions, such as changes in pixels, or from numerical changes in physical simulators, and there are still few models that can predict based on both visual and physical characteristics, like humans. Therefore, in this study, we first propose a model that can predict the collision situation of objects based on both visual information and physical characteristics in the environment, and explain the situation at that time in language. Then, when physical common sense about the environment was added, it was possible to explain the collision situation of objects in as much detail as a human being. In the experiment, we verified the accuracy of change point extraction for each proposed model, the accuracy of sentence generation for the predicted content, and the accuracy of sentence generation after adding common sense. The results of the experiment showed that when the accuracy of the base prediction model is high, the accuracy of extracting change points such as object collisions and the accuracy of text generation that describes the situation at that time also increases. In addition, the sentences generated after adding physical common sense were generated with high accuracy as sentences that take into account the situation of the environment.