高级检索

    基于常识引导与多模态特征融合的视觉关系推理

    Visual Relationship Reasoning via Commonsense Guide and Multimodal Feature Fusion

    • 摘要: 场景图生成作为连接低级视觉感知和高级语义推理的桥梁,在各种视觉语言任务中扮演着重要角色。然而,现有的专用场景图生成模型常因缺乏语义约束机制而产生不合理的关系预测,存在逻辑幻觉。目前,尽管多模态大模型(MLLMs)具备丰富的常识,但直接将其应用于场景图生成任务时,其固有的黑盒特性与严重的幻觉问题导致关系推理结果缺乏可靠性与可解释性,且其参数规模巨大、计算开销高昂,难以满足实际部署需求。为此,本文提出了KETR(Commonsense Knowledge Embedding Transformer),一种嵌入常识知识的端到端场景图生成模型。该模型通过多尺度视觉特征融合编码器提取图像全局与局部特征,利用基于可学习查询的并行解码器实现实体检测。核心创新在于提出软对齐知识融合机制,通过视觉特征与外部常识知识图谱嵌入空间的概率映射动态聚合知识信息,同时引入常识感知的主宾语指示器与知识驱动评分约束,将TransE 模型的评分函数作为正则化项融入联合训练损失,有助于提升关系推理的语义一致性。在Visual Genome 和Open Images V6 两个主流基准数据集上的广泛实验表明,KETR 在多项核心指标上取得了大幅度的性能提升。进一步的消融实验表明,软对齐知识融合有助于改善低频关系的识别能力,常识感知主宾语指示器和知识驱动评分约束能够减少与实体角色、物体可供性及空间语义明显冲突的不合理关系预测。

       

      Abstract: Efficient Scene Graph Generation (SGG) serves as a bridge connecting low-level visual perception and high-level semantic reasoning, playing an important role in various vision-language tasks. However, existing dedicated SGG models often produce unreasonable relation predictions due to the lack of semantic constraint mechanisms. Although multimodal large language models (MLLMs) possess rich commonsense knowledge, their inherent black-box characteristics and severe hallucination issues lead to unreliable relationship reasoning results when directly applied to SGG tasks. To address these challenges, this paper proposes KETR (Commonsense Knowledge Embedding Transformer), an end-to-end scene graph generation model that incorporates commonsense knowledge embeddings. The model extracts global and local image features through a multi-scale visual feature fusion encoder and employs a parallel decoder based on learnable queries for entity detection. The core innovation lies in a soft alignment knowledge fusion mechanism that dynamically aggregates knowledge information through probabilistic mapping between visual features and external commonsense knowledge graph embedding spaces, and a commonsense-aware subject-object indicator and knowledge-driven scoring constraints are introduced, integrating the TransE scoring function as a regularization term into the joint training loss to significantly improve the logical consistency of relationship reasoning. Extensive experiments on Visual Genome and Open Images V6 demonstrate that KETR achieves competitive performance on multiple standard SGG metrics. Further ablation studies show that soft-aligned knowledge fusion improves the recognition of infrequent predicates, while the commonsense-aware subject-object indicator and knowledge-driven scoring constraint reduce relation predictions that conflict with entity roles, object affordances, and spatial semantics.

       

    /

    返回文章
    返回