高级检索

    Web信息检索研究进展

    STATE OF THE ART OF INFORMATION RETRIEVAL ON THE WEB

    • 摘要: Web上大量、分布、动态的信息造成了“信息过载”,如何在传统信息检索技术的基础上开展针对 Web的检索工作已经成为一项重要的研究课题 .但是 ,繁多的 Web信息检索系统和各种模糊的概念给用户的选择和研究人员的讨论带来了不便 .同时 ,有关 Web信息检索最新技术的比较完整的分析又十分缺乏 .在此 ,对 Web信息检索技术进行了综述 ,从 Web信息检索系统的层次化分类 (搜索引擎与目录、元搜索引擎、信息检索 agent)、一般机制和关键新技术 (基于超链的相关度排序、检索结果的联机聚类、基于概念的检索、相关度反馈 )等方面加以阐述 ,以期对感兴趣的同行有参考作用

       

      Abstract: A mass of distributed and dynamic information on the Web has resulted in “information overload”. With the flood of information, it has become an important research issue to search the Web based on traditional information retrieval technology. However, various systems and ambiguous terminology of information retrieval on the Web bring much trouble to users in application and researchers in development as well. Moreover, there are few publications which roundly analyze technologies up to the minute in this field. The state of the art of information retrieval on the Web is surveyed in this paper. First categorized are the systems of sorts in a hierarchical taxonomy and each layer is examined thoroughly. Then analyzed are the general mechanism and several critical new technologies involved in information retrieval on the Web, including results ranking based on hyperlinks, online clustering of retrieval results, concept search, relevance feedback and so on. The paper ends with the authors’ achievement and future work.

       

    /

    返回文章
    返回