<p>Data visualization (DV) is widely used to communicate insights from large datasets. To lower the barrier to DV use, researchers have investigated automatic DV tasks like natural language question (NLQ) to visualization translation, formally called <i>text-to-vis</i>. However, text-to-vis assumes well-organized NLQs expressed in single sentences, which does not reflect real-world scenarios where complex DVs often require <i>consecutive exchanges</i> between users and systems. We introduce Conversational text-to-Visualization (CoVis), a novel task aimed at constructing DVs through multi-turn user-system interactions. We reconstruct and significantly extend Dial-NVBench through an enhanced multi-stage LLM-driven workflow with backwards planning, mandatory SQL execution validation, and intelligent error correction. The reconstructed dataset features improved executability and comprises dialogue sessions with diverse queries−ranging from dataset inquiries and data manipulation to visualization generation−reflecting real-world data visualization complexity. We investigate both neural-based and LLM-based approaches for this task. For the neural approach, we present comprehensive evaluations of MMCoVisNet, a multi-modal network which comprehends dialogue context and employs adaptive decoders: a text decoder for general responses, an SQL-form decoder for data queries (extended to support multi-table joins), and a DV-form decoder for visualizations. For the LLM-based approach, we develop CoVisLLM with a novel query decomposition and reconstruction framework featuring five systematically designed tasks that break down complex reasoning into structured intermediate representations. Comparative evaluation shows improvements over strong baselines for both MMCoVisNet and CoVisLLM.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CoVis: Neural and LLM-Driven Multi-Turn Interactions for Conversational Text-to-Visualization Generation

  • Yuanfeng Song,
  • Jinwei Lu,
  • Raymond Chi-Wing Wong

摘要

Data visualization (DV) is widely used to communicate insights from large datasets. To lower the barrier to DV use, researchers have investigated automatic DV tasks like natural language question (NLQ) to visualization translation, formally called text-to-vis. However, text-to-vis assumes well-organized NLQs expressed in single sentences, which does not reflect real-world scenarios where complex DVs often require consecutive exchanges between users and systems. We introduce Conversational text-to-Visualization (CoVis), a novel task aimed at constructing DVs through multi-turn user-system interactions. We reconstruct and significantly extend Dial-NVBench through an enhanced multi-stage LLM-driven workflow with backwards planning, mandatory SQL execution validation, and intelligent error correction. The reconstructed dataset features improved executability and comprises dialogue sessions with diverse queries−ranging from dataset inquiries and data manipulation to visualization generation−reflecting real-world data visualization complexity. We investigate both neural-based and LLM-based approaches for this task. For the neural approach, we present comprehensive evaluations of MMCoVisNet, a multi-modal network which comprehends dialogue context and employs adaptive decoders: a text decoder for general responses, an SQL-form decoder for data queries (extended to support multi-table joins), and a DV-form decoder for visualizations. For the LLM-based approach, we develop CoVisLLM with a novel query decomposition and reconstruction framework featuring five systematically designed tasks that break down complex reasoning into structured intermediate representations. Comparative evaluation shows improvements over strong baselines for both MMCoVisNet and CoVisLLM.