高级检索

    用数据采掘方法获取汉语词性标注规则

    A DATA MINING METHOD TO ACQUIRE PART OF SPEECH RULES IN CHINESE TEXT

    • 摘要: 从数据采掘的角度对汉语文本词性标注规则的获取进行研究 .在满足用户规定的支持度向量的前提下 ,先从候选集模式中挑选出常用模式 ;然后采掘出具有高可信度的产生式规则 .该过程完全是自动的 ,而获取的规则在表达上是明确的 ,同时又是隐含在数据中的、用户不易发现的 .实验表明 :在原有统计方法的基础上 ,利用自动获得的标注规则作为补充 ,可以提高词性标注的正确率 .

       

      Abstract: A data mining method to acquire part of speech rules in Chinese text is presented. Given an array of support degree, it selects frequent pattern from candidate pattern set. Then it extracts a set of production rules that have high confidence degree. The process is automatic. The rules acquired are clear, but implicit in data set and previously unknown by users. The experiment shows a system that incorporates statistic method with rule method has better performance.

       

    /

    返回文章
    返回