AuDiffusion: multi-agent controlled text-to-image generation with attention-enhanced mamba blocks
摘要
We present AuDiffusion, a diffusion framework that introduces a multi-agent design to improve controllability, semantic alignment, and efficiency in text-to-image generation. The system comprises three cooperating agents responsible for enriching textual input, selecting suitable structural constraints, and performing image synthesis with an enhanced diffusion backbone. This modular design provides more explicit structural control and adaptive decision-making compared with conventional monolithic pipelines, while retaining strong global context modeling. We evaluate AuDiffusion on the ImageNet 256