基于复合结构的知识库分类体系匹配方法

林海伦; 贾岩涛; 王元卓; 靳小龙; 程学旗; 王伟平

doi:10.7544/issn1000-1239.2017.20150843

基于复合结构的知识库分类体系匹配方法

A Composite Structure Based Method for Knowledge Base Taxonomy Matching

摘要

摘要: 近年来，分类体系匹配由于其在知识库构建和融合等方面的广泛应用，已成为国内外工业界和学术界的研究热点.然而，随着网络大数据的不断发展，分类体系变得越来越庞大和复杂，构造一种通用有效的分类体系匹配器以适应大规模、异构分类体系匹配的扩展性仍然面临很大的挑战.为此，提出了一种基于复合结构的分类体系匹配方法BiMWM，该方法利用分类体系中分类的复合结构信息：微观结构和宏观结构，将分类体系匹配问题转化为二部图上的优化问题进行求解.首先，创建赋权的二部图建模分类体系之间候选的匹配类对关系；然后，通过计算二部图上的最大权匹配剪枝选择最优的分类体系的匹配类对.BiMWM方法可以在多项式时间内为2个分类体系产生最优匹配.实验结果表明:与当前先进的基准方法相比，该方法能够有效提升大规模、异构分类体系匹配的性能.

Abstract: Taxonomy matching, i.e., an operation of taxonomy merging across different knowledge bases, which aims to align common elements between taxonomies, has been extensively studied in recent years due to its wide applications in knowledge base population and proliferation. However, with the continuous development of network big data, taxonomies are becoming larger and more complex, and covering different domains. Therefore, to pose an effective and general matching strategy covering cross-domain or large-scale taxonomies is still a considerable challenge. In this paper, we presents a composite structure based matching method, named BiMWM, which exploits the composite structure information of class in taxonomy, including not only the micro-structure but also the macro-structure. BiMWM models the taxonomy matching problem as an optimization problem on a bipartite graph. It works in two stages: it firstly creates a weighted bipartite graph to model the candidate matched classes pairs between two taxonomies, then performs a maximum weight matching algorithm to generate an optimal matching for two taxonomies in a global manner. BiMWM runs in polynomial time to generate an optimal matching for two taxonomies. Experimental results show that our method outperforms the state-of-the-art baseline methods, and performs good adaptability in different domains and scales of taxonomies.

HTML全文

参考文献(0)

施引文献

资源附件(0)