Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Lambda Architecture is a data-processing pattern that runs two paths over incoming data: a batch path recomputes results from historical data, while a speed path processes recent events for fresher results. A serving layer makes results from both paths available to queries. It is designed to combine broad historical processing with low-latency updates, at the cost of operating two processing paths.

How Lambda Architecture works

The pattern divides data processing into three layers. The batch and speed layers handle data on different timelines; the serving layer exposes their results.

Batch layer

The batch layer stores or reads historical data and computes batch views from it. AWS describes a reference design that appends records to an immutable, append-only master dataset and processes that data in the batch path. Reprocessing the historical dataset lets this path produce results across the full available history.

Speed layer

The speed layer processes new or recent events incrementally, so queries can reflect changes before the next batch computation catches up. A CMU-hosted technical chapter describes stream processing as incrementally updating results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Serving layer

The serving layer makes computed views available to query systems. In AWS’s reference diagram, the batch and stream paths feed a merged serving layer for downstream analytics.

Example: transaction totals by region

Imagine a system that answers queries about transaction totals by region. The batch path can periodically calculate totals across historical transactions. Meanwhile, the speed path can process recent transactions and update results before the next batch run. A query service can then return a result that combines the historical view with fresher updates. This is an illustrative example from a CMU-hosted technical chapter, not a claim about a specific deployed system.

When Lambda Architecture may fit

Consider the pattern when a workload needs both historical recomputation and event-driven updates that are fresher than a batch cycle. Its two paths serve complementary needs: one processes historical data broadly, and the other updates results incrementally.

There is no universal data-volume, latency, or cost threshold established for choosing Lambda Architecture. The decision depends on the workload’s requirements and whether the team can build and maintain both paths, along with a coherent way to serve their results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tradeoffs to consider

  • Two paths to operate: Teams must maintain batch and stream processing logic, as well as make their outputs work together for queries.
  • Coordination and reconciliation: Because separate paths can produce results on different schedules, the system needs a coherent serving experience across those outputs.
  • Event-driven concerns are implementation-dependent: AWS notes that event-driven architectures can experience variable latency from network communication and are often eventually consistent. It also describes challenges with transaction handling, duplicates, and determining overall state. These are general event-driven architecture concerns, not automatic properties of every Lambda Architecture implementation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Are particular technologies required?

No. Lambda Architecture is a pattern, not a required product stack. An AWS white paper gives examples in its AWS context, including Amazon EMR and Athena for analytics; Amazon Kinesis Data Streams, Kinesis Data Firehose, and Kinesis Data Analytics for stream or real-time processing; Spark Streaming and Spark SQL on EMR; and Amazon S3 as persistent object storage. These are examples from that reference paper, not requirements or a current recommendation.

For optional deeper reading, Manning’s Big Data: Principles and Best Practices of Scalable Realtime Data Systems includes material on the speed layer and discussion of Kafka and Storm: Manning book page.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.