Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Spark DAG is an execution graph, but the Spark UI shows more than one kind. The Jobs and Stages views map RDD or DataFrame lineage and execution stages; the SQL tab maps query operators and data flow. To understand what happened, follow the graph into its stage and task details, then read shuffle, input/output, and timing metrics alongside it. A diagram alone does not establish a performance root cause.

What a Spark DAG shows

In the Jobs view, a job’s DAG shows a broad processing lineage: vertices represent RDDs or DataFrames, and edges represent operations. Its job details also list stages, their states and task progress, input and output, and shuffle reads and writes. Use this view to trace the broad flow of work and identify where data movement or execution becomes significant.

A stage detail page has a separate DAG visualization. It groups operations into scopes and can display labels such as BatchScan, WholeStageCodegen, and Exchange. For DataFrame and SQL workloads, the stage view can be cross-referenced with the corresponding SQL entry.

These related diagrams are not all the same kind of plan. Jobs and Stages emphasize lineage and execution; the SQL graph emphasizes query operators and their data flow. Apache Spark’s Spark 4.2.0 Web UI documentation describes the current UI views and their metrics. If you use Spark 3.5.6, consult its version-specific Web UI documentation, since UI details can change between releases.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How a job, stage, and task fit together

A Spark job is associated with an action, such as save or collect. The scheduler divides the job into stages, and launches tasks to execute the work in a stage. The graph helps show relationships among operations; stage and task details show how work was scheduled and progressed.

Timing also depends on scheduling. Spark uses FIFO scheduling by default within an application, while fair sharing can be configured. Concurrent jobs and the scheduling mode can therefore affect when work receives resources and how timing appears.

Read a Spark DAG step by step

  1. Open the job. In the Spark UI, go to Jobs and select the relevant job. Note its status, duration, event timeline, associated SQL query if shown, and stage list.
  2. Inspect the stages. Open the relevant stage details. Compare input and output with shuffle read and write, then examine task duration and, where available, scheduler delay, remote shuffle reads, fetch wait, and spill.
  3. Follow SQL work into the SQL view. For a DataFrame or SQL workload, use the associated SQL entry. Read the operator flow and its inline metrics. Expand the plan details when you need the parsed, analyzed, or optimized logical plan, or the physical plan.
  4. Connect evidence before changing code or configuration. Check whether stage and task metrics support what the operator graph suggests. For example, substantial shuffle activity shows data movement, but does not by itself prove which join or setting caused it.

Choose the right view for the question

View What nodes represent Best question to ask Evidence to inspect
Jobs DAG RDDs or DataFrames and the operations connecting them How does the job’s processing flow fit together? Job timeline, stage list, input/output, and shuffle totals
Stage DAG Operation scopes within a stage What work is grouped into this stage, and where should I inspect execution? Task progress and task-level timing and shuffle metrics
SQL graph and plan details Query operators connected by data flow How was the query represented and planned, and what do its operators report? Inline operator metrics and parsed, analyzed, optimized, and physical plans

Interpret metrics without overclaiming

Metrics describe different kinds of work and waiting. Input and output indicate data read or produced; shuffle read and write describe data exchanged between stages. Task duration is elapsed task time, not a direct measure of compute alone. Spark defines scheduler delay as time waiting to be scheduled and shuffle fetch wait as time blocked waiting for shuffle data.

  • Large shuffle reads or writes: evidence of substantial data movement. Inspect the stage and SQL operator context before attributing it to a particular operation.
  • High scheduler delay: time spent waiting to be scheduled, not time executing the task’s computation.
  • High shuffle fetch wait: time blocked waiting for shuffle data; interpret it alongside remote shuffle reads and the stage’s other activity.
  • Spill: a useful indicator to examine with the task and stage context, rather than a diagnosis by itself.

Do not assume each operation in a graph maps one-to-one to a task, or that graph shape or elapsed duration alone explains performance. Use the UI’s actual task and stage status, timelines, and metric definitions to support a diagnosis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Inspect completed applications

The live Spark UI is available only while the application is running. To inspect an application after it ends, configure event logging and use the Spark History Server, which can reconstruct an equivalent UI from persisted application events. See the Apache Spark monitoring documentation for event logging and History Server details.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.