Abstract:
Text classification can help users selectively process huge volumes of texts in the Internet. Text title classification based on example texts is presented in this paper. It not only considers the direct matches between titles and the keyword sets of classes, but also takes into account the upper concept matches and semantic similarities. It uses vector space model as the representation for texts. It adopts the mechanism of indirect matches (upper concept matches), and calculates the similarities between texts and classes in a semantic space rather than term’s space. As a result, it makes full use of the context ofKeywords instead of their frequencies, to determine the degree of correlation between keywords and classes.