BKCrawler: A Scalable Web Data Extraction System Using Weak Supervision
摘要
Automated data collection from diverse websites poses a significant challenge in web mining. This paper introduces BKCrawler, a novel system that automatically detects and extracts information from multiple websites without predefined structure definitions. We propose an integrated architecture comprising two key components: a Site Scouter for automatic website analysis, and a Data Harvester incorporating machine learning models for data collection and extraction. Notably, we apply weak supervision techniques in data labeling for the information extraction model, substantially reducing manual labeling costs while maintaining high accuracy. Experiments on real estate data from 17 websites demonstrate the system’s effectiveness, achieving high accuracy in page classification (F1-score 0.93) and competitive performance in extracting key information fields (F1-scores ranging from 0.65 to 0.88). BKCrawler has been successfully deployed and is currently in use, proving the effectiveness and scalability of the proposed method. This research opens avenues for developing intelligent data extraction systems applicable to various domains.