Advanced Search
    Wu Xuhan, Sun Shiliang. Visual Relationship Reasoning via Commonsense Guide and Multimodal Feature FusionJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660459
    Citation: Wu Xuhan, Sun Shiliang. Visual Relationship Reasoning via Commonsense Guide and Multimodal Feature FusionJ. Journal of Computer Research and Development. DOI: 10.7544/issn1000-1239.202660459

    Visual Relationship Reasoning via Commonsense Guide and Multimodal Feature Fusion

    • Efficient Scene Graph Generation (SGG) serves as a bridge connecting low-level visual perception and high-level semantic reasoning, playing an important role in various vision-language tasks. However, existing dedicated SGG models often produce unreasonable relation predictions due to the lack of semantic constraint mechanisms. Although multimodal large language models (MLLMs) possess rich commonsense knowledge, their inherent black-box characteristics and severe hallucination issues lead to unreliable relationship reasoning results when directly applied to SGG tasks. To address these challenges, this paper proposes KETR (Commonsense Knowledge Embedding Transformer), an end-to-end scene graph generation model that incorporates commonsense knowledge embeddings. The model extracts global and local image features through a multi-scale visual feature fusion encoder and employs a parallel decoder based on learnable queries for entity detection. The core innovation lies in a soft alignment knowledge fusion mechanism that dynamically aggregates knowledge information through probabilistic mapping between visual features and external commonsense knowledge graph embedding spaces, and a commonsense-aware subject-object indicator and knowledge-driven scoring constraints are introduced, integrating the TransE scoring function as a regularization term into the joint training loss to significantly improve the logical consistency of relationship reasoning. Extensive experiments on Visual Genome and Open Images V6 demonstrate that KETR achieves competitive performance on multiple standard SGG metrics. Further ablation studies show that soft-aligned knowledge fusion improves the recognition of infrequent predicates, while the commonsense-aware subject-object indicator and knowledge-driven scoring constraint reduce relation predictions that conflict with entity roles, object affordances, and spatial semantics.
    • loading

    Catalog

      Turn off MathJax
      Article Contents

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return