高级检索

    基于非监督训练的汉语词性标注的实验与分析

    EXPERIMENTS AND ANALYSES OF CHINESE PART OF SPEECH TAGGING BASED ON NON SUPERVISION TRAINING

    • 摘要: 概率参数的获取是基于统计的词性标注的两个主要研究方向之一 .侧重于研究非监督方式 ,利用未标注的语料进行训练获取概率参数 .实现了一个非监督的训练标注模式—— HMM Basic;从不同的初始模型和训练集出发对汉语词性标注进行了实验 ;分析了训练集规模、初始模型的选择对系统标注性能的影响并讨论了其中所存在的问题

       

      Abstract: Probability parameter obtaining is one of the two main study directions of part of speech tagging based on statistics. In this paper, emphasis is laid on non supervision model study, and probability parameters are obtained by training using untagging corpus. A non supervision training tagging model——HMM Basic is implemented. Experiments on Chinese part of speech tagging from different initial models and training sets are made, and the influence on the tagging performance as a result of the selections of the training set size and the initial model is discussed. And the existent problems are also analysed.

       

    /

    返回文章
    返回