ISSN 1000-1239 CN 11-1777/TP

计算机研究与发展 ›› 2016, Vol. 53 ›› Issue (11): 2542-2555.doi: 10.7544/issn1000-1239.2016.20150906

• 人工智能 • 上一篇    下一篇



  1. 1(武汉大学计算机学院 武汉 430072); 2(华东交通大学软件学院 南昌 330013); 3(贵州师范大学大数据与计算机科学学院 贵阳 550001); 4(百度在线网络技术(北京)有限公司 北京 100085) (
  • 出版日期: 2016-11-01
  • 基金资助: 
    国家自然科学基金项目(61133012);国家社会科学基金重大招标项目(11&ZD189);教育部人文社科基金项目(16YJAZH029);江西省科技厅科技攻关项目(20121BBG70050,20142BBG70011);江西省高校人文社科基金项目(XW1502,TQ1503);江西省普通本科高校中青年教师发展计划访问学者专项资金;江西省社科规划项目(16TQ02) This work was supported by the National Natural Science Foundation of China (61133012), the National Social Science Major Tender Project (11&ZD189), the Humanity and Social Science Foundation of Ministry of Education (16YJAZH029), the Science and Technology Research Project of Jiangxi Provincial Department of Science and Technology (20121BBG70050,20142BBG70011), the Humanity and Social Science Foundation of Jiangxi Provincial Universities (XW1502,TQ1503), the Visiting Scholar Special Fund for the Development Plan of Young and Middle-Aged Teachers of General Universities in Jiangxi Province, and the Social Science Planning Project of Jiangxi Province (16TQ02).

Caption Generation from Product Image Based on Tag Refinement and Syntactic Tree

Zhang Hongbin1,2, Ji Donghong1, Yin Lan1,3, Ren Yafeng1, Niu Zhengyu4   

  1. 1(Computer School, Wuhan University, Wuhan 430072); 2(School of Software, East China Jiaotong University, Nanchang 330013); 3(School of Big Data and Computer Science, Guizhou Normal University, Guiyang 550001); 4(Baidu Online Network Technology (Beijing) Co, Ltd, Beijing 100085)
  • Online: 2016-11-01

摘要: 商品图像句子标注是图像标注中一项既有趣又富有挑战的研究任务.噪声单词干扰和句法结构错误是该项研究的制约因素,针对噪声单词干扰,提出关键词精化思想:用绝对排序特征强化关键词权重,完成第1次关键词精化;计算单词的语义相关度评分,进一步优选能准确刻画图像内容的单词,完成第2次关键词精化.设计词序列"拼积木"算法,把关键词拼装成N元词序列.针对句法结构错误,提出句法树思想:基于N元词序列和句法子树递归地构建一棵完整的句法树,遍历该树叶子结点输出句子,标注商品图像.实验结果表明:关键词精化和句法树均有助于改善标注性能,句中的语义信息兼容性和句法模式兼容性得以保持,句子内容更连贯、流畅.

关键词: 图像标注, 商品图像, 句子标注, 关键词精化, 句法树, 词序列“拼积木”, N元词序列, 自然语言生成

Abstract: Automatic caption generation from product image is an interesting and challenging research task of image annotation. However, noisy words interference and inaccurate syntactic structures are the key problems that affect the research heavily. For the first problem, a novel idea of tag refinement (TR) is presented: absolute rank (AR) feature is applied to strengthen the key words weights. The process is called the first tag refinement. The semantic correlation score of each word is calculated in turn and the words that have the tightest semantic correlations with images content are summarized for caption generation. The process is called the second tag refinement. A novel natural language generation (NLG) algorithm named word sequence blocks building (WSBB) is designed accordingly to generate N gram word sequences. For the second problem, a novel idea of syntactic tree (ST) is presented: a complete syntactic tree is constructed recursively based on the N gram word sequences and predefined syntactic subtrees. Finally, sentence is generated by traversing all leaf nodes of the syntactic tree. Experimental results show both the tag refinement and the syntactic tree help to improve the annotation performance. More importantly, not only the semantic information compatibility but also the syntactic mode compatibility of the generated sentence is better retained simultaneously. Moreover, the sentence contains abundant semantic information as well as coherent syntactic structure.

Key words: image annotation, product image, caption generation, tag refinement (TR), syntactic tree (ST), word sequence blocks building, N gram word sequence, natural language generation (NLG)