pandas groupby aggregate element-wise list addition

Question

I have a pandas dataframe that looks as follows:

X                      Y
71455  [334.0,  319.0,  298.0,  323.0]
71455  [3.0,  8.0,  13.0,  10.0]
57674  [54.0,  114.0,  124.0,  103.0]

I want to perform an aggregate groupby that adds the lists stored in the Y columns element-wise. Code I have tried:

df.groupby('X').agg({'Y' : sum})

The result is the following:

                                                   Y
X                                                       
71455  [334.0,  319.0,  298.0,  323.0, 75.0,  55.0,  ...

So it has concatenated the lists and not sum them element-wise. The expected result however is:

X                      Y
71455  [337.0,  327.0,  311.0,  333.0]
57674  [54.0,  114.0,  124.0,  103.0]

I have tried different methods, but could not get this to work as expected.

score 9 · Accepted Answer · answered Aug 20 '18 at 08:41

9

It's possible to use apply on the grouped dataframe to get element-wise addition:

df.groupby('X')['Y'].apply(lambda x: [sum(y) for y in zip(*x)])

Which results in a pandas series object:

X
57674     [54.0, 114.0, 124.0, 103.0]
71455    [337.0, 327.0, 311.0, 333.0]

answered Aug 20 '18 at 08:41

Shaido

27,497
23
70
73

2

Note `apply` is inefficient, effectively a loop. Not recommended with Pandas. – jpp Aug 20 '18 at 09:29

score 5 · Answer 2 · answered Aug 20 '18 at 08:35

Pandas isn't designed for use with series of lists. Such an attempt forces Pandas to use object dtype series which cannot be manipulated in a vectorised fashion. Instead, you can split your series of lists into numeric series before aggregating:

import pandas as pd

df = pd.DataFrame({'X': [71455, 71455, 57674],
                   'Y': [[334.0, 319.0, 298.0, 323.0],
                         [3.0, 8.0, 13.0, 10.0],
                         [54.0, 114.0, 124.0, 103.0]]})

df = df.join(pd.DataFrame(df.pop('Y').values.tolist()))

res = df.groupby('X').sum().reset_index()

print(res)

       X      0      1      2      3
0  57674   54.0  114.0  124.0  103.0
1  71455  337.0  327.0  311.0  333.0

df.pop('Y').values.tolist() can be replaced by list (df.pop('Y')) — Arcanist Lupus, Aug 20 '18 at 17:24
@ArcanistLupus, Yes, it *can* be, but is it necessarily better? I believe NumPy array -> list directly is pretty efficient. — jpp, Aug 20 '18 at 18:08

score 5 · Answer 3 · answered Aug 20 '18 at 08:40

5

If you convert your lists to numpy arrays sum will work:

df['Y'] = df['Y'].apply(np.array)

df.groupby('X')['Y'].apply(np.sum)

#X
#57674     [54.0, 114.0, 124.0, 103.0]
#71455    [337.0, 327.0, 311.0, 333.0]
#Name: Y, dtype: object

answered Aug 20 '18 at 08:40

zipa

27,316
6
40
58

Note `apply` is inefficient, effectively a loop. Not recommended with Pandas. – jpp Aug 20 '18 at 09:29

pandas groupby aggregate element-wise list addition

3 Answers3