From unimodal to multimodal: a framework for generating high-quality multimodal emotional chit-chat dialogue
摘要
Large language models (LLMs) have demonstrated formidable capabilities in both task-oriented and chit-chat dialogues. However, when extended to large vision-language models (LVLMs), we found LVLMs excel in objectively describing the content in the image, while subjective multimodal emotional chit-chat (MEC) dialogues ability is insufficient, a shortfall we attribute to the scarcity of high-quality MEC data. The collection and annotation of high-quality MEC dialogue data are helpful, but the cost of such processes makes the acquisition of large-scale data challenging. Addressing this gap, we introduce an adversarial LLM-based data augmentation framework U2MEC for generating MEC dialogue data from unimodal data. Our framework is inspired by user behaviors in multimodal dialogues, including two patterns: image to prove text and text to explain image, which makes the generation process conformable to human habits. Utilizing U2MEC, we have generated and constructed a Chinese instruction dataset named MEC-chat, which contains 23k high-quality MEC dialogues. We also propose an evaluation benchmark named MECBench to assess the chit-chat abilities of LVLMs. Empirical results confirm that fine-tuning LVLMs on the MEC-chat leads to significant improvements in both MECBench and human assessments.