How to convert a JavaPairRDD to Dataset?

半城伤御伤魂 提交于 2019-12-08 04:13:40

问题


SparkSession.createDataset() only allows List, RDD, or Seq - but it doesn't support JavaPairRDD.

So if I have a JavaPairRDD<String, User> that I want to create a Dataset from, would a viable workround for the SparkSession.createDataset() limitation to create a wrapper UserMap class that contains two fields: String and User.

Then do spark.createDataset(userMap, Encoders.bean(UserMap.class));?


回答1:


If you can convert the JavaPairRDD to List<Tuple2<K, V>> then you can use createDataset method which takes List. See below sample code.

JavaPairRDD<String, User> pairRDD = ...;
Dataset<Row> df = spark.createDataset(pairRDD.collect(), Encoders.tuple(Encoders.STRING(),Encoders.bean(User.class))).toDF("key","value");

or you can convert to RDD

Dataset<Row> df = spark.createDataset(JavaPairRDD.toRDD(pairRDD), Encoders.tuple(Encoders.STRING(),Encoders.bean(User.class))).toDF("key","value");


来源:https://stackoverflow.com/questions/42405905/how-to-convert-a-javapairrdd-to-dataset

易学教程内所有资源均来自网络或用户发布的内容,如有违反法律规定的内容欢迎反馈
该文章没有解决你所遇到的问题?点击提问,说说你的问题,让更多的人一起探讨吧!