A Model for Collecting and Processing Topical Information in the Web and Its Application
-
-
Abstract
Internet is increasingly becoming an important media for news reporting. Presented in this paper is a model for collecting and processing topical information in the web. Selection of sample space, extraction of topical features, and issues in page gathering are described, as well as post processing the data collected. With this model, it is possible to obtain, for a specific topic, the strength of presence in the Internet from a variety of angles, having worth for social scientific research. Based on this model, a system is implemented and an experiment is conducted using "16th Congress" as the topic from October 22nd to November 24th, 2002. It is concluded from the experiment data that the amount of information that are related to the 16th Congress is 7.3% among all the information, and the amount of topical information exhibits a strong taking off from November 2nd, and reached its peak on November 20th, etc.
-
-