Slow performance of POS tagging. Can I do some kind of pre-warming?

nltk, python

Solution

Those first 18 seconds are the POS tagger being unpickled from disk into RAM. If you want to get around this, load the tagger yourself outside of a request function.

import nltk.data, nltk.tag
tagger = nltk.data.load(nltk.tag._POS_TAGGER)

And then replace `nltk.pos_tag` with `tagger.tag`. The tradeoff is that app startup will now take +18seconds.

Problem

I am using NLTK to POS-tag hundereds of tweets in a web request. As you know, Django instantiates a request handler for each request. I noticed this: for a request (~200 tweets), the first tweet needs ~18 seconds to tag, while all subsequent tweets need ~120 milliseconds to tag. What can I do to speed up the process? Can I do a "pre-warming request" so that the module data is already loaded for each request? ``` class MyRequestHandler(BaseHandler): def read(self, request): #this runs for a GET request #...in a loop: tokens = nltk.word_tokenize( tweet) tagged = nltk.pos_tag( tokens) ```

Original source