Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a new Apache Spark streaming application, choose Structured Streaming. Apache Spark describes the older Spark Streaming API—also called DStreams—as a legacy project that no longer receives updates, and recommends Structured Streaming for new applications. The difference is more than a newer API: DStreams model continuous data as a sequence of RDDs, while Structured Streaming expresses streaming work as DataFrame or Dataset queries using Spark SQL.

How do Spark Streaming and Structured Streaming differ?

Spark Streaming and Structured Streaming process continuous data, but they expose different programming models. Apache Spark’s FAQ identifies Spark Streaming as the previous-generation engine and Structured Streaming as the current generation. The Spark overview likewise distinguishes DStreams from the newer DataFrame and Dataset APIs.

Aspect Spark Streaming (DStreams) Structured Streaming
Status Legacy project; Apache Spark says it is no longer updated. Current streaming API recommended by Apache Spark for new applications.
Programming model A continuous stream represented as a sequence of RDDs, transformed with RDD operations. Streaming computations expressed as DataFrame or Dataset queries through Spark SQL.
Time and state Capabilities depend on the DStream application and its operations; the cited Spark guidance does not provide a direct feature-by-feature comparison. Documents event-time windows, watermarks for late data, and state cleanup.
Performance comparison Not established as categorically faster or slower by the cited sources. No like-for-like benchmark against DStreams is established by the cited sources.

How Structured Streaming’s model works

Structured Streaming treats incoming records conceptually as rows appended to a table. You define a query much as you would for a static table, and Spark incrementally updates the query’s result as new data arrives. It does not keep the entire input table in memory; it retains the intermediate state needed to update the result. See the Structured Streaming programming guide for the documented model.

Event time, windows, and late data

Event time is the timestamp recorded in a record. It can differ from processing time—the time Spark receives or handles the record. Event-time windows group records according to when events occurred, rather than when they arrived. A watermark sets a threshold for how late data may be while also allowing Spark to discard old aggregation state. Choose the threshold with the behavior of your data and query in mind: a watermark is part of managing late records and state, not a promise to accept every late event.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does exactly-once processing require?

Exactly-once is conditional, not a blanket guarantee for every pipeline. Spark’s programming guide describes progress tracking with source offsets and checkpoints (and write-ahead logs where applicable). End-to-end exactly-once behavior depends on the source being replayable and the sink being idempotent, so that work can be replayed after a failure without producing duplicate or inconsistent output. Check the semantics of the actual source and sink you use; do not infer an end-to-end guarantee from the API choice alone.

Should you migrate an existing DStream application?

Apache Spark recommends migrating existing Spark Streaming applications to Structured Streaming, but migration is a workload- and version-specific engineering task, not an automatic API swap. Consult the migration guide for the Spark releases involved, and test the behavior of stateful operations, source progress, checkpoints, and output writes before moving production traffic.

Review checkpoints and state

A checkpoint stores progress and, for stateful queries, state needed to resume execution. Some Structured Streaming query settings are coupled to state in a checkpoint. The documentation notes that changing state-partitioning-related settings may require discarding the existing checkpoint and starting a new query. Treat that as a state and recovery decision: understand what data may need replay or rebuilding before replacing a checkpoint, and verify the applicable behavior for the Spark version you deploy.

Check Kafka offset retention

For Kafka, Structured Streaming tracks offsets internally. If Kafka has deleted offsets the query still needs—for example, because of retention—the query can encounter data loss. The Kafka integration guide documents the failOnDataLoss option, which can make the query fail visibly in such cases. Starting-offset settings apply when creating a new query; when resuming an existing query, Spark uses the progress recorded for that query. Confirm that Kafka retention and recovery procedures match the outage and replay scenarios your application must handle.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is one API faster?

The cited Apache Spark documentation does not establish a controlled, like-for-like performance result showing that either API is universally faster. Throughput and latency depend on the workload, Spark version, source and sink, state size, trigger configuration, and cluster setup. If performance is the reason for considering migration, benchmark representative work under the same conditions and measure the outcomes your application cares about rather than relying on a general speed claim.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Practical choice

  • Starting a new Spark streaming pipeline: Use Structured Streaming, following Apache Spark’s current recommendation.
  • Maintaining a DStream pipeline: Plan migration based on your Spark version, stateful logic, checkpoint requirements, and source and sink behavior.
  • Comparing performance: Test your own equivalent workload; the API comparison alone does not establish a winner.

For broader API positioning, see Apache Spark’s Structured Streaming overview.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.