Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apache Ozone connects to Hadoop and Spark in two main ways: use its Hadoop filesystem connector with ofs:// for direct filesystem access, or use the S3A connector with s3a:// when applications already target S3-compatible storage. The first route requires Ozone’s filesystem client JAR and configuration; the second requires an available Ozone S3 Gateway, a matching hadoop-aws dependency, endpoint settings, and credentials.

Choose the connection method

Route Best fit Application path Core requirement
Ozone filesystem connector Hadoop or Spark applications that can use Ozone’s Hadoop-compatible filesystem and need a rooted view across volumes and buckets. ofs://<om-service-id>/<volume>/<bucket>/path/to/key Ozone filesystem client JAR and filesystem configuration.
Ozone S3 Gateway with Hadoop S3A Applications already written for S3-compatible storage or using Hadoop’s S3A connector. s3a://<bucket>/path/to/key Reachable S3 Gateway, compatible hadoop-aws dependency, endpoint configuration, and credentials.

The ofs:// route uses Hadoop-compatible filesystem access. Ozone’s o3fs:// scheme is limited to one bucket; Ozone documentation retains it for legacy compatibility and recommends ofs:// for new deployments. See the Ozone OFS documentation.

S3A can suit existing S3-oriented applications: the Ozone S3 Gateway exposes an S3-compatible REST interface, and Hadoop S3A presents S3 operations through Hadoop’s filesystem interface. The Ozone S3A guide names Hive, Impala, and Spark among tools that can use this route. That does not establish universal S3 feature parity or identical behavior for every application.

Connect through Ozone’s ofs:// filesystem

Configure Hadoop clients

Make the ozone-filesystem-hadoop3 client JAR available on each client’s Hadoop classpath. Configure the filesystem implementation as org.apache.hadoop.fs.ozone.RootedOzoneFileSystem. When Ozone should be the default Hadoop filesystem, the Ozone documentation also shows setting fs.defaultFS to an Ozone Manager URI.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A file URI follows this form: ofs://<om-service-id>/<volume>/<bucket>/path/to/key. Replace the placeholders with the Manager service ID and the Ozone volume and bucket in your environment. The OFS setup guide documents the scheme and configuration.

Make the connector available to Spark

Spark accesses ofs:// through Hadoop-compatible filesystem APIs; it is not a separate Ozone-specific Spark API. Ensure the Ozone client JAR and required filesystem configuration reach the driver and executors. If Spark does not pick up the implementation from its Hadoop configuration, set:

spark.hadoop.fs.ofs.impl=org.apache.hadoop.fs.ozone.RootedOzoneFileSystem

Spark’s regular read and write APIs can then use an ofs:// URI as a data path. The Apache Ozone Spark integration guide provides examples using standard Spark APIs; use the actual Ozone Manager service ID and bucket paths for your deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Connect through S3A and the Ozone S3 Gateway

This route keeps an application’s s3a:// paths, but it depends on the gateway and Hadoop AWS libraries being present and correctly configured. The essential client settings described in the Ozone S3A guide are:

  • Add hadoop-aws to the client and use the same version as hadoop-common.
  • Set fs.s3a.endpoint to the Ozone S3 Gateway endpoint reachable from the client.
  • Set fs.s3a.endpoint.region to a valid-looking logical region; the guide’s example is us-east-1.
  • Set fs.s3a.path.style.access=true, because the gateway uses path-style URLs.
  • Provide the Ozone S3 access and secret keys, or use the documented AWS environment variables.
  • Consider fs.s3a.bucket.probe=0 and fs.s3a.change.detection.mode=none, which the guide identifies as compatibility settings for Ozone.

With security enabled, the guide says to obtain a key and secret with ozone s3 getsecret while authenticated with Kerberos. Treat these as Ozone S3 credentials and follow your cluster’s security procedures for storing and distributing them.

Once configured, S3A can support application reads and writes as well as data movement. Ozone documents Hadoop filesystem commands, copies between local storage and Ozone, and DistCp transfers between HDFS and Ozone as usage patterns in its S3A guide.

Check Spark, Hadoop, and Ozone compatibility

Do not assume a client setup that works in one version combination will work unchanged in another. Apache Ozone’s Spark integration guide says its examples were tested with Spark 3.5.x and Ozone 2.2.0. It also flags a specific runtime issue: Ozone 2.1.0 and later require Hadoop 3.4.x classes for several classes removed from Ozone’s bundled copies, while Spark 3.5.x ships with Hadoop 3.3.4.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the exact Hadoop distribution, Spark build, and Ozone release before rolling out a connector. Follow the applicable release guidance for resolving class availability and version conflicts; do not casually replace Spark’s bundled Hadoop libraries, since the documentation does not establish that such a change is safe across deployments.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Handle authentication and deployment

Kerberos-enabled Spark on YARN

Spark must be allowed to obtain Ozone delegation tokens. For YARN, the guide shows this setting:

spark.kerberos.access.hadoopFileSystems=ofs://ozone1/

The submitting user also needs a valid Kerberos ticket. Use the Ozone Manager service ID for the target environment in the URI rather than copying the example literally. Consult the Spark integration guide for the security configuration.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Spark on Kubernetes

Make the Ozone filesystem client JAR, any required Hadoop compatibility JAR, and core-site.xml available to both driver and executors. Ozone’s guide recommends placing these in a custom Spark image. Its example configuration includes fs.ofs.impl and ozone.om.address; substitute your Manager address, versions, image location, and security settings. Verify that the running pods can load the JARs and configuration, not merely that they are present in the image.

Validate the connection before a cluster rollout

  • Confirm whether the application will use ofs:// or s3a://, and verify the matching connector and URI scheme.
  • Check that the required client JARs and configuration are available on every process that performs filesystem operations, including Spark executors.
  • Verify the endpoint or Ozone Manager service ID, and confirm network access from the application runtime.
  • For S3A, check matching Hadoop AWS and Hadoop Common versions, path-style access, credentials, and the gateway endpoint.
  • For Kerberos, verify the submitting user’s ticket and Spark’s permission to request Ozone delegation tokens.
  • Test a small read and write with the actual runtime and authentication mode before moving a full workload.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.