高级检索

    模型互联网中基于语句置信度的大小模型并行协作评估

    Sentence Confidence-Guided Parallel Collaboration with Large and Small Models for Response Evaluation in AI-Model Network

    • 摘要: 模型互联网作为连接异构模型的新兴计算范式,对其输出质量进行评估是维持系统稳定的关键环节。然而,人工评估成本高昂,传统自动评估难以捕捉深层语义,而新兴的大语言模型裁判策略面临大模型高消耗与小模型低精度的两极分化困境。尽管已有研究尝试通过模型协作来缓解这一矛盾,但现有协作方法多采用串行协作评估模式,存在资源等待问题。为此,提出了基于语句置信度的大小模型并行协作评估方法。构建了“小模型初评-置信度筛选-大模型仲裁”的协作机制,利用小评估模型输出的逻辑值量化评估置信度,精准识别低置信度样本并触发大模型修正;为支撑该协作机制的高效运行,进一步设计了推理与仲裁过程解耦的并行协作评估方法,有效消除了串行机制的等待时延。实验结果表明,在保持与单一大模型同等评估准确率的前提下,能够有效减少约50%的大模型调用,并降低了系统响应时延。实现了评估精度、资源消耗与响应时延的协同优化,为自动化评估技术在模型互联网中的落地应用提供了重要参考。

       

      Abstract: As an emerging paradigm connecting heterogeneous models, the “AI-model network” relies on rigorous output quality assessment to maintain its stability. However, manual evaluation is cost-prohibitive, while traditional methods difficult to capture deep semantics. Furthermore, emerging LLM-as-a-Judge strategies face a dilemma: large models entail high resource consumption, whereas small models suffer from low accuracy. Although collaborative methods have been proposed to mitigate this issue, current approaches rely on serial mechanisms, leading to latency problems due to resource waiting. To address these challenges, this paper proposes a parallel collaborative evaluation method based on sentence confidence. We construct a dynamic workflow comprising “Small Model Initial Evaluation, Confidence Filtering, and Large Model Arbitration”. By utilizing the logit output by the small model to quantify confidence, the method precisely identifies low-confidence samples to trigger correction by the large model. To support efficient operation, we further design a parallel framework that decouples the inference and arbitration processes, effectively eliminating the waiting latency inherent in serial mechanisms. Experimental results indicate that the approach maintains accuracy comparable to a single large model while saving approximately 50% of large model calls and 2reducing system latency. It achieves a coordinated optimization of evaluation precision, resource consumption, and response latency, providing valuable insights for automated evaluation in AI-model network.

       

    /

    返回文章
    返回