Abstract:
The traditional approach for training speaker codebooks only uses one session training speech samples, but the recognition system based on this approach is usually not robust. To adapt to the intraspeaker variations, the paper here introduces an approach for training speaker codebooks using multiple session training speech samples,with every speaker having multiple codebooks. These codebooks are trained based on the minimum recognition error rate.To compensate for the variations arising from transmission conditions, an approach to compensation of the variation presented. To speed up recognition speed, an on line feature extraction method for voiced sounds and two level vector quantization and codebook index strategy are used. These techniques increase the robustness of the speech feature and speed up the training and identification procedure greatly. Finally, the identification results of comparison using the perceptually based linear predictive(PLP) analysis and the LPC cepstrum analysis are given.