高级检索

    基于义类同现频率的汉语语义排歧方法

    A CHINESE SENSE DISAMBIGUATION METHOD BASED ON SENSE CO OCCURRENCE FREQUENCY

    • 摘要: 义类标注是信息检索和自然语言处理中的一个重要问题.但依靠人工对义类进行标注不仅是一个十分烦琐的工作,而且很难把握标准.因此,对义类代码自动标注的研究就显得尤为迫切,而要实现自动标注,必须解决多义词排歧这一重要问题.在对《现代汉语词典》(以下简称《词典》)的义类标注过程中,文中通过统计相邻词语义类组合串的出现频率构造了一个同现频率矩阵集.这一同现频率矩阵集充分利用了义类体系的层次结构,极大地减少了数据稀疏和数据冗余.在此基础上,对《词典》中的多义词进行了排歧,结果较为满意.

       

      Abstract: Sense tagging is very important for the information retrieval and natural language processing. Tagging the sense code by hand is a time consuming work and it is hard to make the sence codes consistent with each other. Therefore the research on automatic sense tagging is a very important work. To do this, the disambiguation of polysemous words must be solved first. In the research of sense tagging in 《Modern Chinese Dictionary》(MCD) , a sense co occurrence frequency matrix (SCFM) set is constructed by making statistics on the frequencies of sense code combination (SCC) of adjacent words that occur in a sample set . This SCFM makes full use of the hierarchy structure of the semantic system, and reduces data sparsness and redundancy. The disambiguation of those polysemous words are done and the words in MCD are tagged by employing a method based on the SCFM set. And the result is satisfying.

       

    /

    返回文章
    返回