Linux: Large int array: mmap vs seek file?

filesystems, linux, memory, memory-management, x86-64

Solution

I'd say performance should be similar if access is truly random. The OS will use a similar caching strategy whether the data page is mapped from a file or the file data is simply cached without an association to RAM.

Assuming cache is ineffective:

- You can use `fadvise` to declare your access pattern in advance and disable readahead.

- Due to address space layout randomization, there might not be a contiguous block of 4 TB in your virtual address space.

- If your data set ever expands, the address space issue might become more pressing.

So I'd go with explicit reads.

Problem

Suppose I have a dataset that is an array of 1e12 32-bit ints (4 TB) stored in a file on a 4TB HDD ext4 filesystem.. Consider that the data is most likely random (or at least seems random). ``` // pseudo-code for (long long i = 0; i < (1LL << 40); i++) SetFileIntAt(i) = GetRandInt(); ``` Further, consider that I wish to read individual int elements in an unpredictable order and that the algorithm runs indefinately (it is on-going). ``` // pseudo-code while (true) UseInt(GetFileInt(GetRand(1<<40))); ``` We are on Linux x86_64, gcc. You can assume system has 4GB of RAM (ie 1000x less than dataset) The following are two ways to architect access: (A) mmap the file to a 4TB block of memory, and access it as an int array (B) open(2) the file and use seek(2) and read(2) to read the ints. Out of A and B which will have the better performance?, and why? Is there another design that will give better performance than either A or B?

Original source