Abstract:
As an emerging paradigm connecting heterogeneous models, the “AI-model network” relies on rigorous output quality assessment to maintain its stability. However, manual evaluation is cost-prohibitive, while traditional methods difficult to capture deep semantics. Furthermore, emerging LLM-as-a-Judge strategies face a dilemma: large models entail high resource consumption, whereas small models suffer from low accuracy. Although collaborative methods have been proposed to mitigate this issue, current approaches rely on serial mechanisms, leading to latency problems due to resource waiting. To address these challenges, this paper proposes a parallel collaborative evaluation method based on sentence confidence. We construct a dynamic workflow comprising “Small Model Initial Evaluation, Confidence Filtering, and Large Model Arbitration”. By utilizing the logit output by the small model to quantify confidence, the method precisely identifies low-confidence samples to trigger correction by the large model. To support efficient operation, we further design a parallel framework that decouples the inference and arbitration processes, effectively eliminating the waiting latency inherent in serial mechanisms. Experimental results indicate that the approach maintains accuracy comparable to a single large model while saving approximately 50% of large model calls and 2reducing system latency. It achieves a coordinated optimization of evaluation precision, resource consumption, and response latency, providing valuable insights for automated evaluation in AI-model network.