find occurrences of a string in a large file that cannot fit memory

bigdata, python

Solution

If you use generators you can access a big file and do the processing.

simple grep command,

def command(f):
    def g(filenames, **kwa):
        lines = readfiles(filenames)
        lines = (outline for line in lines for outline in f(line, **kwa))
        # lines = (line for line in lines if line is not None)
        printlines(lines)
    return g

def readfiles(filenames):
    for f in filenames:
        for line in open(f):
            yield line


def printlines(lines):
    for line in lines:
            print line.strip("\n")

@command
def grep(line, pattern):
    if pattern in line:
        yield line


if __name__ == '__main__':
    import sys
    pattern = sys.argv[1]
    filenames = sys.argv[2:]
    grep(filenames, pattern=pattern)

Problem

I was asked to find the count of occurrences of the string "And" in a large file that is 10GB big and there is 1GB RAM. How would I do it efficiently. I answered that we need to read the file in memory chunks of 100MB each and then find the total occurences of "And" in each memory chunk and keep a cumulative count of the string "And". Interviewer was not satisfied with my answer and he told me how does the command grep work in unix. Write a code similar to that in python but I did not know the answer. I will appreciate answer to this question.

Original source

Related problems