高级检索

    基于向量空间模型的有导词义消歧

    SUPERVISED WORD SENSE DISAMBIGUATION BASED ON VECTOR SPACE MODEL

    • 摘要: 词义消歧一直是自然语言理解中的一个关键问题 ,该问题解决的好坏直接关系到自然语言处理中诸多应用问题的效果优劣 .由于自然语言知识表示的困难 ,在手工规则的词义消歧难以达到理想效果的情况下 ,各种有导机器学习方法被应用于词义消歧任务中 .借鉴前人的成果引入信息检索领域中向量空间模型文档词语权重计算技术来解决多义词义项的知识表示问题 ,并提出了上下文位置权重的计算方法 ,给出了一种基于向量空间模型的词义消歧有导机器学习方法 .该方法将多义词的义项和上下文分别映射到向量空间中 ,通过计算多义词上下文向量与义项向量的距离 ,采用 k- NN(k=1)方法来确定上下文向量的义项分类 .在 9个汉语高频多义词的开放和封闭测试中均取得了突出的成绩 (封闭测试平均正确率为 96 .31% ,开放测试平均正确率为 92 .98% ) ,验证了该方法的有效性

       

      Abstract: Word sense disambiguation(WSD) is the key problem in natural language processing because the result of WSD affects seriously many problems in natural language processing and information retrieval. Because of the failure of manpower on WSD, many supervised methods in machine learning were used on this problem. In this paper, a supervised method is proposed to formalize the senses of polysemous word with interesting term weight based on vector space model, then to deal with WSD with k-NN(k=1). The experiments on 9 Chinese polysemous words in both open test and close test with average accuracy 96.31% in close test and 92.98% in open test show that the method in this paper is very good.

       

    /

    返回文章
    返回