Advanced Search
    Chen Wei, Zhuang Fuzhen. Multi-View Enhanced Graph Attention Network for Multimodal Knowledge Graph CompletionJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660438
    Citation: Chen Wei, Zhuang Fuzhen. Multi-View Enhanced Graph Attention Network for Multimodal Knowledge Graph CompletionJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660438

    Multi-View Enhanced Graph Attention Network for Multimodal Knowledge Graph Completion

    • In recent years, multimodal knowledge graphs have introduced auxiliary information such as entity textual descriptions and images, providing richer semantic sources for entity representation learning in knowledge graph completion. However, most existing multimodal knowledge graph completion methods focus on the alignment or fusion of features from different modalities, typically modeling such information as static external attributes of entities, while insufficiently exploiting the structural context formed by local topology. Since entity semantics are derived not only from their own multimodal attributes but also influenced by neighborhood relations, the lack of relation-constrained structural propagation may cause multimodal features to remain at the level of static semantics, making it difficult to form structured representations with reasoning capability. Particularly in scenarios where modal information contains strong noise, directly relying on initial multimodal features or prematurely fusing heterogeneous modalities may weaken the stability and discriminability of entity representations, thereby limiting the ability to predict complex relation patterns and potential facts. To address these issues, this paper proposes a multimodal knowledge graph completion method based on a Multi-View Enhanced Graph Attention Network, namely MEGAT. The proposed method models structural, visual, and textual information as three independent views under a shared graph topology, and constructs a relation-aware graph attention encoder for each view. In this way, relation-constrained neighborhood propagation and structural semantic modeling can be performed separately in different modal spaces. Subsequently, a gated late-fusion mechanism is employed to dynamically integrate multi-view representations, effectively alleviating the semantic interference caused by the direct fusion of heterogeneous modalities. Experimental results show that the proposed method enhances the ability of entity representations to capture graph structural context and complex relation patterns, achieving competitive completion performance on multiple public datasets.
    • loading

    Catalog

      Turn off MathJax
      Article Contents

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return