Efficiently Read last 'n' rows of CSV into DataFrame

csv, numpy, pandas, python

Solution

I don't think pandas offers a way to do this in `read_csv`.

Perhaps the neatest (in one pass) is to use `collections.deque`:

from collections import deque
from StringIO import StringIO

with open(fname, 'r') as f:
    q = deque(f, 2)  # replace 2 with n (lines read at the end)

In [12]: q
Out[12]: deque(['7,8,9\n', '10,11,12'], maxlen=2)
         # these are the last two lines of my csv

In [13]: pd.read_csv(StringIO(''.join(q)), header=None)

Another option worth trying is to get the number of lines in a first pass and then read the file again, skip that number of rows (minus n) using `read_csv`...

Problem

A few methods to do this: - Read the entire CSV and then use `df.tail` - Somehow reverse the file (whats the best way to do this for large files?) and then use `nrows` argument to read - Somehow find the number of rows in the CSV, then use `skiprows` and read required number of rows. - Maybe do chunk read discarding initial chunks (though not sure how this would work) Can it be done in some easier way? If not, which amongst these three should be prefered and why? Possibly related: - Efficiently finding the last line in a text file - Reading parts of ~13000 row CSV file with pandas read_csv and nrows Not directly related: - How to get the last n row of pandas dataframe?

Original source

Related problems