Shuffle a pandas dataframe by groups

Question

My dataframe looks like this

sampleID  col1 col2
   1        1   63
   1        2   23
   1        3   73
   2        1   20
   2        2   94
   2        3   99
   3        1   73
   3        2   56
   3        3   34

I need to shuffle the dataframe keeping same samples together and the order of the col1 must be same as in above dataframe.

So I need it like this

sampleID  col1 col2
   2        1   20
   2        2   94
   2        3   99
   3        1   73
   3        2   56
   3        3   34
   1        1   63
   1        2   23
   1        3   73

How can I do this? If my example is not clear please let me know.

cs95 · Accepted Answer · 2019-11-22T22:05:18.300

26

Assuming you want to shuffle by sampleID. First df.groupby, shuffle (import random first), and then call pd.concat:

import random

groups = [df for _, df in df.groupby('sampleID')]
random.shuffle(groups)

pd.concat(groups).reset_index(drop=True)

   sampleID  col1  col2
0         2     1    20
1         2     2    94
2         2     3    99
3         1     1    63
4         1     2    23
5         1     3    73
6         3     1    73
7         3     2    56
8         3     3    34

You reset the index with df.reset_index(drop=True), but it is an optional step.

edited Nov 22 '19 at 22:05

answered Aug 09 '17 at 08:58

cs95

379,657
97
704
746

Shouldn't it be `np.random.shuffle(groups)` ? – Mar 31 '19 at 13:36
2

@agcala random.shuffle is better suited for shuffling lists of objects (dfs). – cs95 Mar 31 '19 at 16:02
What's the _ in `[df for _, df in df.groupby('sampleID')]`? – A Merii Feb 24 '20 at 13:05
Does `[df for df in df.groupby('sampleID)]` not achieve the same thing? – A Merii Feb 24 '20 at 13:12
2

@AMerii iterating over grouoBy yields a tuple of (index, group). Since we don't need the index we can use the "don't care var" _ to assign it to and do nothing with it. – cs95 Feb 24 '20 at 16:23

score 9 · Answer 2 · answered Aug 23 '20 at 02:37

I found this to be significantly faster than the accepted answer:

ids = df["sampleID"].unique()
random.shuffle(ids)
df = df.set_index("sampleID").loc[ids].reset_index()

for some reason the pd.concat was the bottleneck in my usecase. Regardless this way you avoid the concatenation.

score 0 · Answer 3 · answered Aug 07 '19 at 13:44

Just to add one thing to @cs95 answer. If you want to shuffle by sampleID but you want to have your sampleIDs ordered from 1. So here the sampleID is not that important to keep. Here is a solution where you have just to iterate over the gourped dataframes and change the sampleID.

groups = [df for _, df in df.groupby('doc_id')]

random.shuffle(groups)

for i, df in enumerate(groups):
     df['doc_id'] = i+1

shuffled = pd.concat(groups).reset_index(drop=True)

        doc_id  sent_id  word_id
   0       1        1       20
   1       1        2       94
   2       1        3       99
   3       2        1       63
   4       2        2       23
   5       2        3       73
   6       3        1       73
   7       3        2       56
   8       3        3       34

Shuffle a pandas dataframe by groups

3 Answers3

Linked

Related