scrapy run spider from script

python, python-2.7, scrapy

Solution

It is simple and straightforward :)

Just check the official documentation. I would make there a little change so you could control the spider to run only when you do `python myscript.py` and not every time you just import from it. Just add an `if __name__ == "__main__"`:

import scrapy
from scrapy.crawler import CrawlerProcess

class MySpider(scrapy.Spider):
    # Your spider definition
    pass

if __name__ == "__main__":
    process = CrawlerProcess({
        'USER_AGENT': 'Mozilla/4.0 (compatible; MSIE 7.0; Windows NT 5.1)'
    })

    process.crawl(MySpider)
    process.start() # the script will block here until the crawling is finished

Now save the file as `myscript.py` and run 'python myscript.py`.

Enjoy!

Problem

I want to run my spider from a script rather than a `scrap crawl` I found this page http://doc.scrapy.org/en/latest/topics/practices.html but actually it doesn't say where to put that script. any help please?

Original source

Related problems