高级检索

    统计与规则并举的汉语句法分析模型

    A Chinese Parsing Model Based on Corpus,Rules and Statistics

    • 摘要: 在自然语言分析中,传统的基于规则的方法和近年兴起的基于统计的方法各有利弊,如何把二者有机的结合起来,以提高分析器的处理能力,是当前计算语言学的重要课题。本文采用依存文法,提出了一种基于依存文法的融合语料库、规则方法和统计方法的汉语分析模型CRSP(Corpus,RuleandStatisticsbasedParser)。该模型的特点是将汉语依存文法分析看作是与词性标注过程等价的一个基于统计的标注过程。文中首先介绍了CRSP的设计思想,然后讨论了从标注过的语料中获取知识的方法,叙述了用于词性标注和依存关系标注的统计模型。试验表明这种模型具有很大的优越性。

       

      Abstract: It is of great significance to take advantages of both rule-based approach and statisticsebased approach to improve the performance of parsers. For this purpose. the authors designed and implemented a new model for Chinese dependency parsing, called CRSP(Corpus, Rule and Statistics -based Parser). The originality of CRSP is that dependency parsing is viewed as both statistics-based tagging process and a single statistics based tagging model is successively used for part of speech (POS) tagging and upper dependency relation (UDR) tagging. In this paper, we first give introduction to the design philosophy of CRSP, and then discuss the knowledge acquisition from an analyzed corpus to support the parsing process, and describe the statistic model for both POS tagging and UDR.tagging. The experiments show that the model is successful.

       

    /

    返回文章
    返回