A Comprehensive Survey of Dark Web Crawlers
摘要
The field of cybercrime investigation has undergone significant change as a result of the widespread use of strong encryption algorithms and sophisticated anonymity routing, which presents difficult circumstances for law enforcement agencies (LEAs). Consequently, law enforcement agencies (LEAs) are increasingly relying on unencrypted web information or anonymous communication networks (ACNs) as potential sources of leads and evidence for their investigations. LEAs have access to a significant tool for gathering and storing potentially important data for investigative purposes: automated web content harvesting from servers. Although web crawling has been studied since the early days of the internet, relatively little research has been done on web crawling on the “dark web” or ACNs like IPFS, Freenet, Tor, I2P, and others. This work offers a thorough systematic literature review (SLR) with the goal of investigating the characteristics and prevalence of dark web crawlers. After removing pointless entries, a refined set of 30 peer-reviewed publications about crawling and the dark web remained from an original pool of 30 articles. According to the review, most dark web crawlers are written in Python and frequently use Selenium or Scrapy as their main web scraping libraries. The lessons learned from the SLR were applied to the creation of an advanced Tor-based web crawling model that was easily incorporated into an already existing software toolbox designed for ACN-focused research. It also highlights promising directions for future study in this quickly developing sector, highlighting how crucial it is to use cutting-edge technologies to effectively fight cybercrime in a digital environment that is becoming more complicated.