高级检索

    基于最大熵方法的中英文基本名词短语识别

    Chinese and English BaseNP Recognition Based on a Maximum Entropy Model

    • 摘要: 使用了基于最大熵的方法识别中文基本名词短语 在开放语料ChineseTreeBank上 ,只使用词性标注 ,达到了平均 87 4 3% / 88 0 9%的查全率 /准确率 由于 ,关于中文的基本名词短语识别的结果没有很好的可比性 ,又使用相同的算法 ,尝试了英文的基本名词短语识别 在英文标准语料TREEBANKⅡ上 ,开放测试达到了 93 31% / 93 0 4 %的查全率/准确率 ,极为接近国际最优水平 这既证明了此算法的行之有效 ,又表明该方法的语言无关性

       

      Abstract: A maximum entropy model in Chinese BaseNP recognition is used in this paper The open test on Chinese TreeBank, the public corpus, indicates the average recall and precision of 87 43% and 88 09% respectively with limited knowledge (text itself and its POS tag) Because of the incomparability of Chinese BaseNP recognition results, the same algorithm is applied in English BaseNP recognition The test on TREEBANK Ⅱ shows that the recall and precision are 93 31% and 93 04%, which are close to the state of the art This not only proves the availability of the algorithm, but also indicates its language independence

       

    /

    返回文章
    返回