高级检索

    基于义原同现频率的汉语词义排歧方法

    A CHINESE WORD SENSE DISAMBIGUATION METHOD BASED ON PRIMITIVE CO-OCCURRENCE DATA

    • 摘要: 词义排岐是自然语言处理的重点和难点问题之一 .基于语料库的统计方法已被广泛地应用于词义排岐 .大多数的统计方法都受到数据稀疏的困扰 ,对于词义排岐而言 ,由于有大量同义词的存在 ,数据稀疏问题变得更为严重 .充分利用“知网”这个知识源的特性 ,提出了一种基于义原同现频率的词义排岐方法 ,在很大程度上克服了数据稀疏问题 .此外 ,该方法还避免了繁重的人工标注语料的过程 ,通过在一个约 10万字的语料库上获得义原同现频率矩阵 ,并以此作为词义排岐的依据 .实验表明 ,该方法对词义排岐具有较高的正确率

       

      Abstract: Word sense disambiguation is one of the difficult problems and a key point in natural language processing. Corpus based sense disambiguation methods, like most other statistical NLP approaches, suffer from the problem of data sparseness. Especially, because there are a great number of synonyms in a text, this problem in word sense disambiguation becomes worse. In this paper, an approach is described, which overcomes this problem using the property of the Hownet. Using the word definition in the Hownet, the primitive co occurrence data matrix is obtained, which are collected from a corpus of about 100000 characters without any manual tagging. Finally, this method is tested and the result shows that it has higher accuracy.

       

    /

    返回文章
    返回