Abstract:
Word sense disambiguation is one of the difficult problems and a key point in natural language processing. Corpus based sense disambiguation methods, like most other statistical NLP approaches, suffer from the problem of data sparseness. Especially, because there are a great number of synonyms in a text, this problem in word sense disambiguation becomes worse. In this paper, an approach is described, which overcomes this problem using the property of the Hownet. Using the word definition in the Hownet, the primitive co occurrence data matrix is obtained, which are collected from a corpus of about 100000 characters without any manual tagging. Finally, this method is tested and the result shows that it has higher accuracy.