Abstract:
A complete introduction to the model selection, ad hoc the mixture model, for clustering analysis is included in this paper, and the key related technologies are discussed seriatim, Based on these, the author introduces the Bayesian posteriori model selection, which reduces the complexity of the algorithm based on the mixture model and improves the precision (against the traditional model selection). To estimate the parameters in the posteriori model, two different Bayesian estimation methods, maximum likelihood estimation, and conditional expectation estimation, are compared. The posteriori model based hierarchical clustering algorithms are described, with the analysis of the domain itself. Results of high accuracy have been achieved in experiments for real world text clustering.