高级检索

    面向可解释人机协同的知识引导主动学习方法

    Knowledge-Guided Active Learning for Explainable Human-AI Collaboration

    • 摘要: 机器学习在真实场景中的有效应用通常依赖充足可靠的标注数据,而许多任务的标注需要领域专家参与,成本较高。主动学习通过选择高价值未标注样本交由专家标注,以降低数据构建成本。现有方法多从学习器视角设计查询准则,如不确定性、多样性、委员会分歧和核心集覆盖等,但其跨数据集、模型和预算条件的收益并不稳定,甚至可能劣于随机标注。同时,传统主动学习通常将专家限定为标签提供者,缺少解释样本选择依据和模型预测逻辑的协同媒介。专家反馈因此难以从样本标签上升为可复用的知识规则,模型也难以与专家领域知识有效对齐。针对上述问题,本文提出面向可解释人机协同的知识引导主动学习方法KALE。KALE以一阶逻辑规则作为模型与专家的协同媒介,先将当前模型预测行为归纳为紧凑逻辑规则,再由专家依据领域知识判别规则正确性。对正确逻辑规则,KALE执行演绎式弱标注扩展;对错误逻辑规则,KALE将其视为知识-模型分歧的符号化观测,并据此开展反绎式分歧查询。方法同时结合类别级规则、局部规则和稀疏规则选择,以控制交互规模。实验表明,KALE较已有基线取得稳定性能提升,在小预算标注场景下优势更加明显。

       

      Abstract: Label scarcity remains a major obstacle to deploying machine learning in real applications, because many tasks require costly annotation by domain experts. Active learning reduces this cost by selecting valuable unlabeled instances for expert annotation. Existing methods usually design query criteria from a learner-centric view, such as uncertainty, diversity, committee disagreement, and core-set coverage. However, their gains are not stable across datasets, models, and annotation budgets, and an unsuitable strategy may even underperform random annotation. Meanwhile, conventional active learning often restricts experts to the role of label providers, lacking a collaborative medium that explains why samples are selected and what decision logic the model follows. As a result, expert feedback is difficult to elevate from individual labels into reusable knowledge rules, and the model cannot be effectively aligned with expert domain knowledge. To address these issues, this paper proposes KALE, a knowledge-guided active learning method for explainable human-AI collaboration. KALE uses first-order logic rules as the collaborative medium between the model and experts. It first induces compact logic rules from the current model behavior, and then asks experts to judge their correctness according to domain knowledge. Correct rules are used for deductive weak-label expansion, while incorrect rules are treated as symbolic observations of knowledge-model disagreement and used to drive abductive disagreement querying. KALE further combines class-level rules, local rules, and sparse rule selection to keep expert interaction manageable. Experiments show that KALE achieves more robust improvement than existing active learning baselines, especially under small annotation budgets.

       

    /

    返回文章
    返回