To run Spark on Kubernetes, submit the application in cluster mode with a Kubernetes API server URL, a Spark container image available to the cluster, and an application entry point. Kubernetes schedules the driver pod Spark creates; that driver then creates executor pods. Start with a fixed executor count to validate permissions, image access, networking, and resource requests. Enable dynamic allocation only when workload variability makes it useful, and use shuffle tracking because Kubernetes does not support Spark’s external shuffle service.
How Spark runs on Kubernetes
Spark’s cluster-mode submission starts a driver pod in Kubernetes. The driver coordinates the application and creates executor pods to do the work; Kubernetes schedules those pods onto available nodes. Completed executor pods terminate. The completed driver pod remains until it is garbage-collected or you remove it manually, so account for driver-pod cleanup when operating repeated jobs.
This is a Kubernetes deployment model, not a special Spark API. The same basic workflow applies to conformant managed clusters such as EKS, GKE, and AKS, to self-managed clusters, and to local clusters such as kind or minikube. The cluster must be able to pull the image and reach the required Kubernetes and application services.
Prepare the cluster and image
Make an image the cluster can run
Build or choose a Spark container image containing the runtime and application dependencies your job needs. Publish it to a registry the cluster can access, or otherwise make it available to every node that may run a pod. The driver and executors must be able to use the selected image; a path that exists only on your submission machine is not enough.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Set up driver permissions and DNS
The driver’s Kubernetes service account needs permission to create the pods, services, and ConfigMaps required by the application. Kubernetes DNS must also be configured so the driver and executors can resolve cluster services. Use a dedicated service account with only the required access rather than relying on broad permissions.
Before scaling out, verify that the driver can authenticate to the API server, that image-pull credentials work in the target namespace, and that pods can communicate as expected. Permission denials, image-pull failures, DNS problems, and insufficient node capacity are common reasons an otherwise valid submission does not start cleanly.
Submit a Spark application in cluster mode
Use spark-submit with a k8s:// master URL, --deploy-mode cluster, a name, an application class, an image, and the application resource. This representative command uses Spark’s example application; replace the API server, image, and JAR values with those for your environment.
./bin/spark-submit
--master k8s://https://<k8s-apiserver-host>:<port>
--deploy-mode cluster
--name spark-pi
--class org.apache.spark.examples.SparkPi
--conf spark.executor.instances=5
--conf spark.kubernetes.container.image=<spark-image>
local:///path/to/examples.jar
The five executors are an example setting, not a recommended default or capacity guarantee. The local:/// application path must be valid in the selected Spark image. For an application-specific job, substitute its main class and packaged resource.
Rank #3
Set the namespace and resource profile
Choose the Kubernetes namespace deliberately; set it with spark.kubernetes.namespace when you need to target a namespace other than the configured default. Configure driver and executor CPU and memory for the job. Spark derives Kubernetes pod requests and limits from its driver and executor core, memory, and overhead settings, so check both the resulting pod resources and the capacity available on eligible nodes. A request that exceeds available capacity can leave pods pending rather than make the job run faster.
Begin with a small, known workload and inspect the driver and executor pod events and logs. Increase resources only after confirming that the current run is constrained by capacity or workload demand—not by RBAC, image access, or connectivity.
Choose fixed executors or dynamic allocation
| Approach | Best fit | Trade-off to manage |
|---|---|---|
| Fixed executor count | Predictable workloads, or jobs where steady resource use and a stable executor pool matter. | Executors remain at the configured level even when the workload needs less, while a count set too low can limit parallel work. |
| Dynamic allocation | Workloads whose demand changes enough that adding and removing executors is useful. | Shuffle data can keep executors alive; scale behavior and resource use depend on allocation limits and timeout settings. |
Dynamic allocation is disabled by default. To use it on Kubernetes, explicitly enable both dynamic allocation and shuffle tracking:
--conf spark.dynamicAllocation.enabled=true
--conf spark.dynamicAllocation.shuffleTracking.enabled=true
Set the initial, minimum, and maximum executor counts and idle timeouts to match the workload and the cluster’s capacity. Shuffle tracking allows Spark to retain executors that hold shuffle data; that can prevent useful data from being lost, but it can also delay executor removal and keep resources occupied. Monitor executor counts and idle-timeout behavior under representative jobs before relying on a particular scale-down pattern.
Best Value
Favor fixed allocation when predictability or a latency target outweighs utilization savings. Consider dynamic allocation when demand varies, but account for shuffle retention and contention from other workloads. Neither setting substitutes for Kubernetes capacity planning.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Control placement and sharing
Kubernetes placement and fairness are separate from Spark’s choice of executor count. Use namespace boundaries to organize workloads; node selectors or pod templates to influence where Spark pods run; and priority classes when jobs need differentiated scheduling priority. Make sure selected nodes have the resources and image access the pods require.
For a shared cluster that needs queueing or reservations, a custom scheduler such as Volcano or YuniKorn can provide additional queue and priority behavior. These are operational choices for cluster scheduling, not prerequisites for running Spark. Confirm that the scheduler and the Spark pod configuration agree on how workloads should be placed before applying them to production jobs.
Choose between spark-submit and the Spark Kubernetes Operator
| Operational need | Direct spark-submit |
Spark Kubernetes Operator |
|---|---|---|
| Workflow | Imperative: submit a run with command-line arguments and configuration. | Declarative: define a SparkApplication resource describing the application. |
| Scheduling integration | Configure Spark and Kubernetes scheduling behavior for each job or submission workflow. | Provides scheduling-related fields; queue behavior can also depend on cluster schedulers and their configuration. |
| Monitoring and cleanup | Plan how to observe runs and remove completed driver pods when needed. | Provides monitoring and a timeToLiveSeconds cleanup field for completed applications. |
| Repeatability | Repeatable when commands and configuration are managed consistently. | Application settings are represented in a Kubernetes resource, which suits declarative operations. |
| Operational ownership | Fits teams that want to own submission and lifecycle handling directly. | Requires installing and operating the operator as well as managing application resources. |
Use direct submission for a straightforward job or when a pipeline already owns the submit-and-monitor lifecycle. Choose the operator when the team wants applications represented as Kubernetes resources and wants its monitoring, scheduling, dynamic-allocation, and TTL cleanup fields. The Kubeflow Spark Operator documentation describes its workflow as applicable to conformant Kubernetes clusters rather than a particular cloud or distribution; the operator is an additional component, not a requirement for Spark-on-Kubernetes.
Validate a run before increasing scale
- Confirm the driver pod starts in the intended namespace and uses the expected service account.
- Check pod events for authorization errors, image-pull failures, scheduling delays, and resource shortages.
- Verify driver-to-executor connectivity and DNS resolution for services the job uses.
- Inspect driver and executor logs to distinguish application errors from Kubernetes startup or scheduling problems.
- For dynamic allocation, observe executor growth, shuffle-related retention, and timeout behavior before raising the maximum executor count.
- Decide how completed driver pods will be cleaned up, either through the operator’s TTL field or an explicit operational cleanup process.
Only increase executor counts after these checks pass. More requested executors do not create more cluster capacity, and pending pods will not improve application throughput.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

