Spark Java remote HDFS cluster


What's inside this article ⌄
  • Spark Java remote HDFS access
  • Spark connect external HDFS
  • Spark read remote HDFS Java
  • Spark Java cross cluster HDFS

Remote HDFS Spark Configuration

Suppose we need to work with different HDFS (clusterB, for instance) from our Spark Java application, running on clusterA.

Firstly, you need to add --conf key to your run command. Depends on a Spark version:

(Spark 1.x-2.1) spark.yarn.access.namenodes=hdfs://clusterA,hdfs://clusterB
(Spark 2.2+) spark.yarn.access.hadoopFileSystems=hdfs://clusterA,hdfs://clusterB

Secondly, when you’re creating Spark’s Java context, add that:

javaSparkContext.hadoopConfiguration().addResource(new Path("core-site-clusterB.xml"));
javaSparkContext.hadoopConfiguration().addResource(new Path("hdfs-site-clusterB.xml"));

You need to go to clusterB and gather core-site.xml and hdfs-site.xml from there (default location for Cloudera is /etc/hadoop/conf) and put near your app running in clusterA.

Pay attention to these points:

  • we are specifying both core-site.xml and hdfs-site.xml, not just one of them
  • we are sending Path object to addResource() method, not just ordinary String!

Troubleshooting

If you’re facing java.net.UnknownHostException: clusterB error, then try to put full namenode address of your remote HDFS with port (instead of hdfs/cluster short name) to --conf into your running command:

(Spark 1.x-2.1) spark.yarn.access.namenodes=hdfs://clusterA,hdfs://namenode.fqdn:port
(Spark 2.2+) spark.yarn.access.hadoopFileSystems=hdfs://clusterA,hdfs://namenode.fqdn:port

If you face AccessControlException error, check this article.

If it’s not working for you, or you’re facing with another error, check “Troubleshooting” section here.