Predicting Plain Text Imageability for Faithful Prompt-Conditional Image Generation
摘要
Text imageability is often used to quantize the ease with which a natural language description can invoke a mental image in a reader. With the proliferation of artificial intelligence powered text-to-image generation models, it will likely play an even more significant role in bridging the gap between language and visual representation. Unfortunately, automatically suggesting proper imageable textural prompts from a piece of plain text has scarcely been systematically investigated. In this paper, we narrow the gap by introducing a novel framework for text imageability assessment to automatically predict whether a piece of plain text and a prompt is highly imageable to be fed into a text-to-image model for faithful image generation. We have also developed a new visual-text dataset, named Ted1.6k, to facilitate model training and validation. Experiment results demonstrate the effectiveness of the proposed method in promoting prompt-guided image generation.