ISSN 1000-1239 CN 11-1777/TP

Journal of Computer Research and Development ›› 2016, Vol. 53 ›› Issue (11): 2542-2555.doi: 10.7544/issn1000-1239.2016.20150906

Previous Articles     Next Articles

Caption Generation from Product Image Based on Tag Refinement and Syntactic Tree

Zhang Hongbin1,2, Ji Donghong1, Yin Lan1,3, Ren Yafeng1, Niu Zhengyu4   

  1. 1(Computer School, Wuhan University, Wuhan 430072); 2(School of Software, East China Jiaotong University, Nanchang 330013); 3(School of Big Data and Computer Science, Guizhou Normal University, Guiyang 550001); 4(Baidu Online Network Technology (Beijing) Co, Ltd, Beijing 100085)
  • Online:2016-11-01

Abstract: Automatic caption generation from product image is an interesting and challenging research task of image annotation. However, noisy words interference and inaccurate syntactic structures are the key problems that affect the research heavily. For the first problem, a novel idea of tag refinement (TR) is presented: absolute rank (AR) feature is applied to strengthen the key words weights. The process is called the first tag refinement. The semantic correlation score of each word is calculated in turn and the words that have the tightest semantic correlations with images content are summarized for caption generation. The process is called the second tag refinement. A novel natural language generation (NLG) algorithm named word sequence blocks building (WSBB) is designed accordingly to generate N gram word sequences. For the second problem, a novel idea of syntactic tree (ST) is presented: a complete syntactic tree is constructed recursively based on the N gram word sequences and predefined syntactic subtrees. Finally, sentence is generated by traversing all leaf nodes of the syntactic tree. Experimental results show both the tag refinement and the syntactic tree help to improve the annotation performance. More importantly, not only the semantic information compatibility but also the syntactic mode compatibility of the generated sentence is better retained simultaneously. Moreover, the sentence contains abundant semantic information as well as coherent syntactic structure.

Key words: image annotation, product image, caption generation, tag refinement (TR), syntactic tree (ST), word sequence blocks building, N gram word sequence, natural language generation (NLG)

CLC Number: