Abstract:
Sense tagging is very important for the information retrieval and natural language processing. Tagging the sense code by hand is a time consuming work and it is hard to make the sence codes consistent with each other. Therefore the research on automatic sense tagging is a very important work. To do this, the disambiguation of polysemous words must be solved first. In the research of sense tagging in 《Modern Chinese Dictionary》(MCD) , a sense co occurrence frequency matrix (SCFM) set is constructed by making statistics on the frequencies of sense code combination (SCC) of adjacent words that occur in a sample set . This SCFM makes full use of the hierarchy structure of the semantic system, and reduces data sparsness and redundancy. The disambiguation of those polysemous words are done and the words in MCD are tagged by employing a method based on the SCFM set. And the result is satisfying.