高级检索

    粒度引导多模型协作的生成式图像分类方法

    Generative Image Classification Method via Granularity-Guided Multi-Model Collaboration

    • 摘要: 近年来,大型视觉-语言模型(vision-language models,VLM)在图像分类任务中展现出卓越的零样本预测能力,推动了图像识别技术的发展。现有VLM普遍依赖于一组预定义的类别标签或外部文本知识库作为语义先验,以实现图像与语义之间的对齐与匹配。然而,在实际应用中,预定义标签往往难以覆盖开放世界中的所有概念,导致模型的适应性和泛化能力受到限制。为应对这一挑战,提出了一种更具开放性的任务——生成式图像分类(generative image classification,GIC),该任务不提供任何先验的语义标签,而目标类别潜藏于一个开放的,可能无限扩展的语义概念空间中,极大地增加了分类难度。针对该任务的复杂性,提出了一种基于粒度引导的多模型协作框架,充分发挥VLM在图像理解与文本生成方面的优势。具体而言,该框架通过设计高效的提示模板,引导VLM在不同语义粒度下生成潜在标签;随后,语言模型基于图像描述自动关联同一粒度内的相关类别,扩展候选标签空间;最后,通过在嵌入空间中进行语义相似度检索,筛选出最相关的最终标签。在多个细粒度图像分类基准数据集上进行的实验验证了所提方法在GIC任务中的有效性与先进性。

       

      Abstract: In recent years, large Vision-Language Models (VLM) have demonstrated remarkable zero-shot prediction capabilities in image classification tasks. Existing VLM-based classification approaches typically rely on a set of predefined category labels or external textual knowledge bases as semantic prior to align and match visual content with language. However, in real-world applications, such predefined labels often fail to cover the vast and dynamic concept space of the open world, thereby limiting the adaptability and generalization ability of these models. To address this challenge, a more open-ended task, Generative Image Classification (GIC), is introduced, in which no prior semantic labels are provided, and the target categories reside in an open-ended, potentially unbounded semantic space, substantially increasing the classification difficulty. To tackle the complexity of GIC, a granularity-guided multi-model collaboration framework is proposed, designed to fully leverage the strengths of VLM in both image understanding and text generation. In this framework, prompt templates are carefully crafted to guide VLM in generating candidate labels at varying levels of semantic granularity. Subsequently, the candidate label space is expanded by a language model that associates semantically related categories based on the generated image descriptions at each granularity level. Finally, the most relevant label is identified through a semantic similarity retrieval process in the embedding space. The effectiveness and superiority of the proposed method for the GIC task are validated through extensive experiments on multiple fine-grained image classification benchmarks.

       

    /

    返回文章
    返回