AI-based Text-to-video (T2V) generation aims to generate videos by comprehending text semantics. Although researchers have achieved useful results in this field, challenges remain in generating long videos, including resource intensity and training-inference gaps. Meanwhile, large language model (LLM), known for generating coherent long texts, offers potential solutions. However, the long text generation capacity of LLM cannot guarantee the consistency of the content of long videos. This paper presents Endless Movie Maker, an LLM-based agent framework that enhances T2V generation by ensuring logical coherence and visual consistency across extended video lengths. The framework integrates scriptwriting, storyboarding, and refinement agents to guide video generation while allowing flexibility and expandability. We do extensive experiments and propose a series of evaluation indicators to illustrate the optimal performance of our method in long video generation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Endless Movie Maker: Zero-Shot Agent System for Making Long Video with Text

  • Xin Zheng,
  • Ziqi Ma,
  • Hang Yu

摘要

AI-based Text-to-video (T2V) generation aims to generate videos by comprehending text semantics. Although researchers have achieved useful results in this field, challenges remain in generating long videos, including resource intensity and training-inference gaps. Meanwhile, large language model (LLM), known for generating coherent long texts, offers potential solutions. However, the long text generation capacity of LLM cannot guarantee the consistency of the content of long videos. This paper presents Endless Movie Maker, an LLM-based agent framework that enhances T2V generation by ensuring logical coherence and visual consistency across extended video lengths. The framework integrates scriptwriting, storyboarding, and refinement agents to guide video generation while allowing flexibility and expandability. We do extensive experiments and propose a series of evaluation indicators to illustrate the optimal performance of our method in long video generation.