高级检索

    基于多智能体辩论的图文交织生成方法

    Interleaved Text-Image Generation Method Based on Multi-Agent Debate

    • 摘要: 多模态检索增强生成通过引入外部文本与视觉知识,为复杂文档理解、知识问答和步骤化任务指导等任务场景提供了重要支撑。然而,现有方法通常将图文交织生成过程交由单一多模态大模型统一完成,即把检索得到的文本块、候选图像及其图像描述同时输入同一个模型,并要求模型在一次生成过程中同时承担文本内容组织、视觉证据筛选、图像插入位置判断以及多图顺序规划等任务。但在复杂长上下文场景下,候选内容中常包含弱相关文本、相似图像和冗余视觉证据,容易导致关键图像漏选、无关图像误选及多图顺序错位等问题。针对上述问题,提出一种基于多智能体辩论机制的多模态检索增强图文交织生成方法。该方法将生成过程解耦为文本推理、视觉感知和全局裁决3类协同角色,并通过双轨草稿生成、集合冲突与顺序冲突检测、多轮对抗式辩论及全局裁决与融合机制,对图像选择、动态补图、冗余剔除和顺序重排进行统一优化。实验结果表明,方法在MRAMG-Bench的网络数据、学术文档和生活方式3个领域上均取得较优表现。特别是在生活方式领域,相比当前强基线M2IO-R1,提出方法在图像召回率和顺序一致性指标上分别提升7.02%和7.77%;在网络数据领域,提出方法相比最优完整基线的平均性能提升7.90%。上述结果验证了提出方法在复杂长文档图文交织生成任务中的有效性。为便于结果复现,提供匿名代码仓库与复现说明:https://github.com/mira-ai-lab/MRAMG-MultiAgent-Debate

       

      Abstract: Multimodal retrieval-augmented generation introduces external textual and visual knowledge, providing important support for complex document understanding, knowledge-based question answering, and step-by-step task guidance. However, existing methods usually delegate the interleaved text-image generation process to a single multimodal large language model. Specifically, the retrieved text chunks, candidate images, and image descriptions are simultaneously fed into the same model, which is required to perform text organization, visual evidence selection, image insertion position determination, and multi-image ordering within a single generation process. In complex long-context scenarios, however, the candidate content often contains weakly relevant text, visually similar images, and redundant visual evidence, which can easily lead to missing key images, selecting irrelevant images, and misordering multiple images. To address these problems, we propose a multimodal retrieval-augmented interleaved text-image generation method based on a multi-agent debate mechanism. The proposed method decomposes the generation process into three collaborative roles: a text reasoning agent, a visual perception agent, and a global judge agent. Through dual-track draft generation, set conflict and ordering conflict detection, multi-round adversarial debate, and global judgment and fusion, the method jointly optimizes image selection, dynamic image supplementation, redundant image removal, and image reordering. Experimental results show that the proposed method achieves competitive performance across the network data, academic document, and lifestyle domains of MRAMG-Bench. In particular, in the lifestyle domain, compared with the strong baseline M2IO-R1, the proposed method improves image recall and ordering consistency by 7.02 and 7.77 percentage points, respectively. In the network data domain, the proposed method improves the average performance by 7.90 percentage points compared with the best complete baseline. These results demonstrate the effectiveness of the proposed method in complex long-document interleaved text-image generation tasks. Anonymous code and reproduction instructions are available at: https://github.com/mira-ai-lab/MRAMG-MultiAgent-Debate

       

    /

    返回文章
    返回