高级检索

    EAFCC:基于共性—特性协同建模与多专家融合的多特征图像描述模型

    EAFCC: A Multi-Feature Image Captioning Model with Common-Specific Collaboration and Mixture-of-Experts Fusion

    • 摘要: 针对现有图像描述方法难以同时兼顾全局语义与细粒度细节、生成结果易偏泛化或关键信息遗漏的问题,本文提出一种基于共性—特性协同建模与多专家融合的多特征图像描述生成模型EAFCC(Expert-Aware Fusion with Common-Specific Collaboration)。该模型并行接收网格特征、区域特征与分割特征,在编码阶段首先通过预融合模块对多源异构特征进行隐式对齐。在共享语义空间中提取跨特征一致的共性信息,同时为各特征流保留差异化表达基础。随后利用共享Transformer编码器建模共性表征,并通过对比学习约束进一步增强其一致性与判别性。同时,采用三条独立编码分支对不同特征流进行特性建模,以保留与特征类型相关的细粒度补充信息,并借助多专家门控路由机制实现特性信息的动态加权融合。最后,将增强后的共性特征与融合后的特性特征进行协同整合,经跳跃连接编码器后送入Transformer解码器生成图像描述。在MS-COCO数据集上的实验表明,EAFCC在CIDEr-D指标上取得了较优性能,并优于现有单特征或双特征方法,验证了共性与特性联合建模的有效性。

       

      Abstract: Existing image captioning methods often struggle to simultaneously capture global semantics and fine-grained details, leading to overly generic descriptions or missing key information. To address this issue, we propose EAFCC, short for Expert-Aware Fusion with Common-Specific Collaboration, a multi-feature image captioning model with mixture-of-experts fusion. The model takes grid features, region features, and segmentation features as parallel inputs. In the encoding stage, a pre-fusion module first performs implicit alignment of heterogeneous visual features and models cross-feature common information within a shared semantic space, while preserving the basis for feature-specific representations. A shared Transformer encoder is then employed to learn the common representations, whose consistency and discriminability are further enhanced by a contrastive learning constraint. Meanwhile, three independent encoding branches are used to model feature-specific representations, so as to retain fine-grained complementary information associated with different feature types. A gated routing mechanism inspired by mixture-of-experts is further introduced to dynamically weight and fuse the specific information. Finally, the enhanced common features and the fused specific features are collaboratively integrated, processed by a skip-connected encoder, and fed into a Transformer decoder to generate image captions. Experimental results on the MS-COCO dataset show that EAFCC achieves competitive performance on the CIDEr-D metric and outperforms existing single-feature and dual-feature methods, demonstrating the effectiveness of common-specific collaborative modeling and dynamic multi-feature fusion for image captioning.

       

    /

    返回文章
    返回