高级检索

    多视图增强图注意力网络的多模态知识图谱补全

    Multi-View Enhanced Graph Attention Network for Multimodal Knowledge Graph Completion

    • 摘要: 近年来,多模态知识图谱通过引入实体文本描述、图像等辅助信息,为知识图谱补全中的实体表示学习提供了更加丰富的语义来源。然而,现有多模态知识图谱补全方法大多侧重于不同模态特征的对齐或融合,通常将这些信息作为实体的静态外部属性进行建模,而对局部拓扑共同构成的结构上下文利用不足。由于实体语义不仅来源于自身的多模态属性,也受到其邻域关系的影响,缺乏关系约束下的结构传播容易使多模态特征停留在静态语义层面,难以形成具有推理能力的结构化表示。尤其在模态信息噪声较强的场景下,直接依赖初始多模态特征或过早融合异构模态,容易削弱实体表示的稳定性和判别性,从而限制对复杂关系模式和潜在事实的预测能力。为解决上述问题,本文提出一种基于多视图增强图注意力网络的多模态知识图谱补全方法MEGAT(Multi-View Enhanced Graph Attention Network)。该方法将结构、视觉和文本信息建模为共享图拓扑下的三个独立视图,并为每个视图构建关系感知图注意力编码器,以在不同模态空间中分别执行关系约束下的邻域传播和结构语义建模。随后,采用门控后融合机制动态整合多视图表示,有效缓解了直接融合异构模态带来的语义干扰。实验结果显示,所提出的方法能够增强实体表示对图结构上下文和复杂关系模式的感知能力,在多个公开数据集上实现竞争性补全性能。

       

      Abstract: In recent years, multimodal knowledge graphs have introduced auxiliary information such as entity textual descriptions and images, providing richer semantic sources for entity representation learning in knowledge graph completion. However, most existing multimodal knowledge graph completion methods focus on the alignment or fusion of features from different modalities, typically modeling such information as static external attributes of entities, while insufficiently exploiting the structural context formed by local topology. Since entity semantics are derived not only from their own multimodal attributes but also influenced by neighborhood relations, the lack of relation-constrained structural propagation may cause multimodal features to remain at the level of static semantics, making it difficult to form structured representations with reasoning capability. Particularly in scenarios where modal information contains strong noise, directly relying on initial multimodal features or prematurely fusing heterogeneous modalities may weaken the stability and discriminability of entity representations, thereby limiting the ability to predict complex relation patterns and potential facts. To address these issues, this paper proposes a multimodal knowledge graph completion method based on a Multi-View Enhanced Graph Attention Network, namely MEGAT. The proposed method models structural, visual, and textual information as three independent views under a shared graph topology, and constructs a relation-aware graph attention encoder for each view. In this way, relation-constrained neighborhood propagation and structural semantic modeling can be performed separately in different modal spaces. Subsequently, a gated late-fusion mechanism is employed to dynamically integrate multi-view representations, effectively alleviating the semantic interference caused by the direct fusion of heterogeneous modalities. Experimental results show that the proposed method enhances the ability of entity representations to capture graph structural context and complex relation patterns, achieving competitive completion performance on multiple public datasets.

       

    /

    返回文章
    返回