Converting tuples to multiple indices in a Pandas Dataframe
pandas, python
Solution
If you want to convert index of your dataframe:
>>> df.index = pd.MultiIndex.from_tuples(df.index)
>>> df
bar__sum foo__sum
100000550 ActivityA 14 12
100001799 ActivityB 7 3
>>> df.index.names = ['id', 'act_type']
>>> df
bar__sum foo__sum
id act_type
100000550 ActivityA 14 12
100001799 ActivityB 7 3
You can also create DataFrame directly from dictionary (`d` is your dict, don't call your variable `dict` since it'll shadow standard python dictionary):
>>> pd.DataFrame(d.values(), index=pd.MultiIndex.from_tuples(d.keys(), names=['id', 'act_type']))
bar__sum foo__sum
id act_type
100001799 ActivityB 7 3
100000550 ActivityA 14 12
Note that `values()` and `keys()` are always in the same order, so no worries about that.
Problem
I'm starting with a dictionary like this: ``` dict = {(100000550L, u'ActivityA'): {'bar__sum': 14.0, 'foo__sum': 12.0}, (100001799L, u'ActivityB'): {'bar__sum': 7.0, 'foo__sum': 3.0}} ``` Which, when converted to a DataFrame, puts as column headers the tuples of (id, activitytype): ``` df = DataFrame(dict).transpose() bar__sum foo__sum (100000550, ActivityA) 14 12 (100001799, ActivityB) 7 3 ``` How can I convert the tuples in the index to a MultiIndex? Ie, so that the end result looks like this instead: ``` bar__sum foo__sum id act_type 100000550 ActivityA 14 12 100001799 ActivityB 7 3 ``` What's the best way to do this? Is there some option on the DataFrame creation that I'm missing? Or should it happen via a list comprehension, which feels inefficient to me.