Large scale MicroBlog location data capture method based on dynamic web page parsing
Yu Ji et al.
What the paper says
Due to the large scale of data, the deviation coefficient of the captured data is large and the capture efficiency is low. To this end, a large-scale Weibo location data retrieval method based on dynamic web page parsing is proposed. Firstly, based on the source of Weibo location data, artificial neural models and random functions are introduced to calculate the weights of feature data. Next, generate a feature vector table and classifier model, and filter the feature text using the established classification model. Finally, by matching the feature data of Weibo location data between dynamic script sites and web pages, a dynamic script parsing framework for Weibo location data on web pages is constructed, and dynamic web page parsing technology is used to capture Weibo location data. The experimental results show that the proposed method has only a 0.1% error in data capture bias, and the capture efficiency reaches 99%. Therefore, this method can significantly improve the crawling effect of large-scale Weibo location data and has certain feasibility.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.