
Webhose.io, a company that provides direct access to live data from hundreds of thousands of forums, news and blogs, posted an article describing a tiny, multi-threaded web crawler, written in Python. This Python web crawler is capable of searching the entire web for you.
Ran Geva, the author of this tiny python web crawler, says:
“I wrote about how it can and does download thousands of pages from multiple sites in just a few hours. No set up or external imports required, just run the following Python code with a 'seed site' and sit back (or do some other work, because it will take a few hours, or even days – depending on the amount of data you need).
https://www.secnews.gr/
The Python based multi-threaded crawler is quite simple and very fast. It is able to detect and eliminate duplicate links and save both the source and the link, which you can then use to find inbound / outbound links for calculating page rank.
It's completely free and the code is as you see below:

Enter the above code with a name like this, e.g. “myPythonCrawler.py”.
To start crawling simply type:
[alert variation=”alert-success”]$ python myΡythonCrawler.py https://www.secnews.gr[/alert]
Then sit back and enjoy your python web crawler.
https://www.secnews.gr/101029/the-first-person-hack-body/See also: Meet the first person to hack his own body to…..
Did you find the article interesting? We look forward to your comments below!
