How to convert a DataFrame back to normal RDD in pyspark?

Question

I need to use the

(rdd.)partitionBy(npartitions, custom_partitioner)

method that is not available on the DataFrame. All of the DataFrame methods refer only to DataFrame results. So then how to create an RDD from the DataFrame data?

Note: this is a change (in 1.3.0) from 1.2.0.

Update from the answer from @dpangmao: the method is .rdd. I was interested to understand if (a) it were public and (b) what are the performance implications.

Well (a) is yes and (b) - well you can see here that there are significant perf implications: a new RDD must be created by invoking mapPartitions :

In dataframe.py (note the file name changed as well (was sql.py):

@property
def rdd(self):
    """
    Return the content of the :class:`DataFrame` as an :class:`RDD`
    of :class:`Row` s.
    """
    if not hasattr(self, '_lazy_rdd'):
        jrdd = self._jdf.javaToPython()
        rdd = RDD(jrdd, self.sql_ctx._sc, BatchedSerializer(PickleSerializer()))
        schema = self.schema

        def applySchema(it):
            cls = _create_cls(schema)
            return itertools.imap(cls, it)

        self._lazy_rdd = rdd.mapPartitions(applySchema)

    return self._lazy_rdd

score 106 · Answer 1 · edited Aug 05 '16 at 22:01

106

Use the method .rdd like this:

rdd = df.rdd

edited Aug 05 '16 at 22:01

gsamaras

71,951
46
188
305

answered Mar 18 '15 at 17:36

dapangmao

2,727
3
22
18

1

yes you are correct. I updated the OP after digging deeper into this. – WestCoastProjects Mar 18 '15 at 17:48
21

yes but it convert to org.apache.spark.rdd.RDD[org.apache.spark.sql.Row] but not org.apache.spark.rdd.RDD[string] – Venu A Positive Jan 28 '16 at 13:10
1

Technically, it's a property: https://spark.apache.org/docs/latest/api/python/pyspark.sql.html#pyspark.sql.DataFrame.rdd – flow2k Aug 26 '20 at 11:53

score 86 · Accepted Answer · edited Jun 21 '16 at 03:54

86

@dapangmao's answer works, but it doesn't give the regular spark RDD, it returns a Row object. If you want to have the regular RDD format.

Try this:

rdd = df.rdd.map(tuple)

or

rdd = df.rdd.map(list)

edited Jun 21 '16 at 03:54

Kristian

21,204
19
101
176

answered May 17 '16 at 21:13

kennyut

3,671
28
30

4

This should be the default behaviour imo when calling `df.rdd` – Nov 30 '17 at 13:59
This is probably a more precise answer actually – WestCoastProjects May 14 '18 at 20:49
What is the `df`, how to initialize it? – David Wei Dec 23 '18 at 14:49
@DavidWei some Dataframe instance, so whatever variable your dataframe is assigned to – lampShadesDrifter Jan 18 '20 at 00:49
1

What are tuple and list ? – Itération 122442 Oct 27 '20 at 10:43
`Row` is a subclass of tuple, so in many most cases doing this is superfluous. – Aaron Zolnai-Lucas Apr 25 '23 at 08:36

score 7 · Answer 3 · edited Nov 01 '22 at 20:55

7

Answer given by kennyut/Kistian works very well but to get exact RDD like output when RDD consist of list of attributes e.g. [1,2,3,4] we can use flatmap command as below,

rdd = df.rdd.flatMap(list)

or

rdd = df.rdd.flatMap(lambda x: list(x))

edited Nov 01 '22 at 20:55

ZygD

22,092
39
79
102

answered May 14 '18 at 17:39

Nilesh

71
1
3

1

This looks like a helpful contribution. – WestCoastProjects May 14 '18 at 20:49

How to convert a DataFrame back to normal RDD in pyspark?

3 Answers3

Linked