Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apache Livy lets you interact with an Apache Spark cluster through a REST API. To get started, install Spark separately, point Livy to that installation with SPARK_HOME, configure Hadoop settings if your cluster needs them, start Livy, then create a session or submit a batch job over HTTP.

What Livy does—and what it does not include

Livy is the REST-facing service; Spark is the separate processing runtime. The Apache Livy project overview describes it as a way to interact with a Spark cluster over REST. It supports interactive Scala and Python work and batch submissions in Scala, Java, or Python. Livy does not bundle Spark, so a working Spark installation is a prerequisite.

Check Spark compatibility and configuration

The official Livy getting-started guide specifies Spark 3.0 or higher and Scala 2.12 builds of Spark. The project repository README likewise says Spark 3.0+ and notes that the runtime Spark can be selected through SPARK_HOME without rebuilding Livy. These are version-sensitive requirements: confirm compatibility against the Livy and Spark versions and distribution used by your cluster.

  • SPARK_HOME should point to the Spark installation Livy will use.
  • For the documented local-session example, set HADOOP_CONF_DIR to the Hadoop configuration directory.
  • If Livy should use Spark configuration from a location other than the configuration under SPARK_HOME, set SPARK_CONF_DIR before starting the server.

Use paths appropriate to your operating system and cluster; the documentation examples do not specify universal values for these variables.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install and start Livy

  1. Get a Livy package. Follow the project’s download instructions and unpack or install the package according to its distribution. The getting-started guide does not establish one package-independent installation command.
  2. Install Spark separately. Set SPARK_HOME to the Spark installation. For a local session that requires Hadoop configuration, set HADOOP_CONF_DIR as well.
  3. Set optional Spark configuration. If needed, set SPARK_CONF_DIR before launching Livy.
  4. Start the service. From the Livy installation directory, run ./bin/livy-server start.
  5. Connect to the service. Livy listens on port 8998 by default. Set livy.server.port to use a different port.

The start command and configuration names above are from Livy’s getting-started documentation; the actual environment-variable values depend on your installation and cluster layout.

Make a first REST request

Livy’s REST API reference documents POST /sessions to create an interactive session. The request selects a session kind such as Scala, Python, or R, and can include resource settings and Spark configuration. For example, a minimal Scala session request has this shape:

curl -X POST -H 'Content-Type: application/json' 
  -d '{"kind":"spark"}' 
  http://localhost:8998/sessions

Consult the deployed version’s REST API reference for the accepted kind values and request fields; valid options depend on the Livy and Spark environment. The response identifies the session. Use the API’s session endpoints to inspect its state and submit interactive work.

For a one-off job rather than an interactive context, Livy also exposes batch submission endpoints, with endpoints for batch state and logs. The same API reference documents these operations. Choose an interactive session when you need a continuing shell context; choose batch submission when you want to submit a job without maintaining an interactive session.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose how Spark runs

The getting-started guide strongly recommends running Spark applications in YARN cluster mode. In that arrangement, YARN accounts for user-session resources in the cluster, and the machine hosting Livy is less likely to be overloaded when several sessions run. Local and other deployment arrangements are possible, but their requirements depend on the cluster; the cited setup material does not provide a complete deployment decision matrix.

The REST API includes resource controls such as driver and executor memory and cores, as well as Spark configuration. Check the API documentation and your cluster’s policy before setting them; available values and their validity are environment-dependent.

Rank #4
The SQL Programming Language: .
  • Used Book in Good Condition

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.