Advanced Search
    Ma Jie, Zhang Hanchi, Qu Ning, Xue Haoquan, Wang Xinping, Liu Jun. Interleaved Text-Image Generation Method Based on Multi-Agent DebateJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660338
    Citation: Ma Jie, Zhang Hanchi, Qu Ning, Xue Haoquan, Wang Xinping, Liu Jun. Interleaved Text-Image Generation Method Based on Multi-Agent DebateJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660338

    Interleaved Text-Image Generation Method Based on Multi-Agent Debate

    • Multimodal retrieval-augmented generation introduces external textual and visual knowledge, providing important support for complex document understanding, knowledge-based question answering, and step-by-step task guidance. However, existing methods usually delegate the interleaved text-image generation process to a single multimodal large language model. Specifically, the retrieved text chunks, candidate images, and image descriptions are simultaneously fed into the same model, which is required to perform text organization, visual evidence selection, image insertion position determination, and multi-image ordering within a single generation process. In complex long-context scenarios, however, the candidate content often contains weakly relevant text, visually similar images, and redundant visual evidence, which can easily lead to missing key images, selecting irrelevant images, and misordering multiple images. To address these problems, we propose a multimodal retrieval-augmented interleaved text-image generation method based on a multi-agent debate mechanism. The proposed method decomposes the generation process into three collaborative roles: a text reasoning agent, a visual perception agent, and a global judge agent. Through dual-track draft generation, set conflict and ordering conflict detection, multi-round adversarial debate, and global judgment and fusion, the method jointly optimizes image selection, dynamic image supplementation, redundant image removal, and image reordering. Experimental results show that the proposed method achieves competitive performance across the network data, academic document, and lifestyle domains of MRAMG-Bench. In particular, in the lifestyle domain, compared with the strong baseline M2IO-R1, the proposed method improves image recall and ordering consistency by 7.02 and 7.77 percentage points, respectively. In the network data domain, the proposed method improves the average performance by 7.90 percentage points compared with the best complete baseline. These results demonstrate the effectiveness of the proposed method in complex long-document interleaved text-image generation tasks. Anonymous code and reproduction instructions are available at: https://github.com/mira-ai-lab/MRAMG-MultiAgent-Debate
    • loading

    Catalog

      Turn off MathJax
      Article Contents

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return