Abstract:
To acquire chunks from running texts is useful for many applications, such as machine translation, information retrieving, etc.. Described in this paper are the schemes of rule-based chunker and statistics-based chunker. Also proposed is a method to combine rule-based processing with statistics-based processing. According to the practical situation the mistake recall is introduced to rate the performance of the system. Compared with the rule-based system, the precision and recall are enhanced to identify chunks, and the error rate is reduced about 7%. The performance of the whole system has been improved greatly.