How can I strip the whitespace from Pandas DataFrame headers?

Question

I am parsing data from an Excel file that has extra white space in some of the column headings.

When I check the columns of the resulting dataframe, with df.columns, I see:

Index(['Year', 'Month ', 'Value'])
                     ^
#                    Note the unwanted trailing space on 'Month '

Consequently, I can't do:

df["Month"]

Because it will tell me the column is not found, as I asked for "Month", not "Month ".

My question, then, is how can I strip out the unwanted white space from the column headings?

This [answer](https://stackoverflow.com/a/36082588/7758804) should be accepted, not the current one. — Trenton McKinney, Nov 07 '21 at 12:21

TomAugspurger · Accepted Answer · 2021-10-24T16:33:06.197

194

You can give functions to the rename method. The str.strip() method should do what you want:

In [5]: df
Out[5]: 
   Year  Month   Value
0     1       2      3

[1 rows x 3 columns]

In [6]: df.rename(columns=lambda x: x.strip())
Out[6]: 
   Year  Month  Value
0     1      2      3

[1 rows x 3 columns]

Note: that this returns a DataFrame object and it's shown as output on screen, but the changes are not actually set on your columns. To make the changes, either use this in a method chain or re-assign the df variabe:

df = df.rename(columns=lambda x: x.strip())

edited Oct 24 '21 at 16:33

answered Feb 06 '14 at 15:49

TomAugspurger

28,234
8
86
69

df.rename(columns=lambda x: x.strip()m axis=1) is necessary here so that the lambda fxn iterate through headers rather than index – spencerlou Oct 04 '22 at 19:53

score 111 · Answer 2 · edited Nov 27 '21 at 00:16

Since version 0.16.1 you can just call .str.strip on the columns:

df.columns = df.columns.str.strip()

Here is a small example:

In [5]:
df = pd.DataFrame(columns=['Year', 'Month ', 'Value'])
print(df.columns.tolist())
df.columns = df.columns.str.strip()
df.columns.tolist()

['Year', 'Month ', 'Value']
Out[5]:
['Year', 'Month', 'Value']

Timings

In[26]:
df = pd.DataFrame(columns=[' year', ' month ', ' day', ' asdas ', ' asdas', 'as ', '  sa', ' asdas '])
df
Out[26]: 
Empty DataFrame
Columns: [ year,  month ,  day,  asdas ,  asdas, as ,   sa,  asdas ]


%timeit df.rename(columns=lambda x: x.strip())
%timeit df.columns.str.strip()
1000 loops, best of 3: 293 µs per loop
10000 loops, best of 3: 143 µs per loop

So str.strip is ~2X faster, I expect this to scale better for larger dfs

score 15 · Answer 3 · answered Apr 23 '19 at 14:17

15

If you use CSV format to export from Excel and read as Pandas DataFrame, you can specify:

skipinitialspace=True

when calling pd.read_csv.

From the documentation:

skipinitialspace : bool, default False
Skip spaces after delimiter.

answered Apr 23 '19 at 14:17

Eric Duminil

52,989
9
71
124

1

This doesn't skip trailing spaces per the OP's example. There doesn't seem to be a reasonable way to do this, particularly for multi-row headers which create MultiIndexes. It can be done, but it should be easier. – Terry Brown Jul 20 '21 at 02:27
@TerryBrown: It doesn't help in the general case, that's true, and also why my answer begins with an "if". I've often seen whitespaces in Dataframes imported from CSV, that's why I mentioned it. – Eric Duminil Jul 20 '21 at 06:04

score 4 · Answer 4 · answered Jul 29 '21 at 17:02

4

If you are looking for an unbreakable way to do it, I would suggest:

data_frame.rename(columns=lambda x: x.strip() if isinstance(x, str) else x, inplace=True)

answered Jul 29 '21 at 17:02

loicgasser

1,403
12
17

Upvoted! This is where my mind went since I like to strip whitespace earlier in my process flow and handle incoming data with variable headers (nans, ints, etc). Using the isinstance(var, type) check slows it down sure - but how many headers are we talking? Here I'd exchange the flexibility for computation since I don't forsee bringing in a header set of more than 25 columns...and definitely not more than 500... – jameshollisandrew Oct 28 '21 at 17:27

score 4 · Answer 5 · answered Aug 14 '21 at 01:59

4

Actually can do that with

df.rename(str.strip, axis = 'columns')

Which is shown in Pandas documentation here.

answered Aug 14 '21 at 01:59

Jervine Lovesu

61
6

How can I strip the whitespace from Pandas DataFrame headers?

5 Answers5

Linked

Related