Lambda vs Kappa: Stream Processing Architecture Choices

Philip Rehberger Aug 11, 2026 6 min read

Compare batch+stream vs stream-only data architectures. Operational complexity, replay semantics, and tooling fit.

When data has to be processed both in batches (overnight aggregations, historical reports) and in real time (live dashboards, fraud detection), you have two architectural choices. The Lambda architecture says do both — run a batch pipeline and a streaming pipeline side by side. The Kappa architecture says do everything in the streaming pipeline and call batch a special case of streaming.

These names come from Nathan Marz's original Lambda paper and Jay Kreps's response a few years later. The debate is older than most engineers reading this, but the underlying question — should historical data and live data flow through the same pipeline? — is more relevant than ever as streaming tools mature.

The Two Architectures Side by Side

Lambda architecture has three layers:

            ┌─────────────────┐
Data ──────►│  Batch layer    │──────►┐
            │  (full history) │       │
            └─────────────────┘       │
            ┌─────────────────┐       │   ┌──────────────┐
            │  Speed layer    │──────►│──►│ Serving layer │──► Queries
            │  (recent only)  │       │   └──────────────┘
            └─────────────────┘       │
            ┌─────────────────┐       │
            │  Master dataset │──────►┘
            └─────────────────┘

The batch layer reprocesses everything from scratch on a schedule. The speed layer processes only recent events as they arrive. The serving layer merges results from both — batch for everything older than the speed layer's window, speed for the rest.

Kappa architecture uses one stream-processing pipeline for everything. Historical processing is just rewinding the stream from the beginning.

            ┌────────────────────┐
Data ──────►│ Immutable log      │──► Stream processor ──► Serving layer
            │ (Kafka or similar) │
            └────────────────────┘
                      ▲
                      │
                      └── replay for reprocessing

When the processing logic changes, you spin up a new instance of the stream processor that reads from the beginning of the log. When it catches up, you swap the old version for the new.

Where the Approaches Came From

Lambda was a response to the limitations of stream processors in the early 2010s. Storm and similar frameworks could compute approximate results in real time but were not reliable enough to be the system of record. Batch was reliable; streaming was fast; combining both got you both properties. The cost was running two pipelines that had to compute the same thing the same way.

Kappa emerged once stream processors got exactly-once semantics and durable, replayable logs (Kafka). If your stream is durable and your processor is reliable, you do not need a separate batch pipeline — you replay the stream when you need a historical view.

What Lambda Actually Costs

The architectural elegance of Lambda hides operational pain. You write the same logic twice — once in Spark, once in Flink (or whatever). When a bug is in the streaming layer, you have to fix it in the batch layer too, and reconcile. The two layers drift apart, and reconciliation becomes its own job.

The operational tax is high enough that most teams who tried Lambda eventually moved off it. The split-implementation problem outweighs the "best tool for each job" argument.

What Kappa Actually Costs

Kappa is simpler, but it depends on infrastructure assumptions:

  • The event log must be durable for as long as the longest reprocessing window. Years of events at scale gets expensive.
  • The stream processor must support reading from arbitrary offsets at high throughput.
  • State management has to handle reprocessing — running the new processor in parallel with the old until catch-up.
  • Reprocessing time matters. If reprocessing the whole stream takes a week, "deploy a new version" becomes a week-long operation.

Most modern stream processors (Flink, Spark Structured Streaming) handle most of this. Kafka handles the log. The pieces exist; the operational sophistication to use them well is what is in short supply.

Reprocessing in Practice

The Kappa reprocessing workflow is the part teams underestimate:

1. Deploy new processor version (v2), consuming from offset 0
2. v2 writes results to a new output topic / table / index
3. Monitor v2 catch-up: how close to head is it?
4. When v2 is at head, switch reads from v1's output to v2's
5. Decommission v1 and its output

Doing this without dropping any events, without doubling cost forever, and without breaking downstream consumers is a real engineering effort. Tools like Flink's savepoints make state transfer feasible but not free.

A Realistic Middle Ground

In practice, most production systems are neither pure Lambda nor pure Kappa. They are stream-first with batch reinforcement:

  • A streaming pipeline processes events as they arrive and feeds a serving layer
  • A batch job runs daily to recompute aggregates that are expensive to maintain in stream
  • The batch results overwrite the streaming results for the windows the batch covers
  • For everything newer than the batch window, the streaming results stand

This is not Lambda — there is one processing technology, one source of truth. It is also not pure Kappa — batch exists. Calling it "Lambda-flavored Kappa" is more accurate than either pure name.

When to Pick Which

Choose Kappa-style stream-first when:

  • Real-time results are the primary requirement
  • Your stream processor supports reliable state and replay
  • The team has streaming experience
  • Event volumes are manageable to retain for reprocessing

Choose batch-first with stream overlays when:

  • Historical accuracy matters more than real-time latency
  • Reprocessing is expensive enough that you only want to do it deliberately
  • Most of your queries hit data older than a few minutes
  • The team is more comfortable with batch tooling

Choose Lambda only when:

  • You inherited it. New systems should not pick Lambda in 2026.

Tooling

Layer Common Choices
Event log Kafka, Pulsar, Kinesis
Stream processor Flink, Spark Structured Streaming, Kafka Streams
Batch (when used) Spark, dbt, Airflow-orchestrated SQL
Serving ClickHouse, Druid, Pinot, Postgres

The serving layer is the underrated choice. It determines what queries are fast, what data freshness you can offer, and how much you spend per query. Picking the serving layer well makes both Lambda and Kappa easier; picking it poorly makes both painful.

The Bigger Point

The Lambda vs Kappa debate is, in 2026, less about which architecture wins and more about how much you trust streaming infrastructure. As that trust has grown, Kappa-style systems have become the default. Lambda is mostly a legacy reality, not a design choice.

If you are designing a new system, the question is not "Lambda or Kappa" but "what serving layer do my queries actually need, and what streaming pipeline can I afford to operate well?" Get those right and the architectural label takes care of itself.


Designing a data pipeline and wondering whether real-time is worth the operational cost? We help teams pick architectures that match the queries their business actually runs. scopeforged.com

Share this article

Related Articles

Need help with your project?

Let's discuss how we can help you build reliable software.