Data Engineering (Spark & Notebooks) · Lakehouse · Streaming & Real-Time

Data architecture patterns: what actually competes with what

Lambda, Kappa, Lakehouse, Mesh, Medallion. They get set out side by side as though the job is to pick one. They answer different questions, and a real platform runs several at once.


The short version

  • Lambda, Kappa, Lakehouse, Mesh, Fabric and Medallion get set out side by side, as though the job is to choose between them. It isn't.
  • Lambda and Kappa answer how data is processed. Lake, Warehouse and Lakehouse answer where it lives. Mesh and Fabric answer who owns it and how it is governed. Those are different questions, so the patterns are not alternatives to one another.
  • A working enterprise platform commonly runs several of them together. The useful skill is knowing which question each one answers, not memorising definitions.
  • The genuine either/or choices are narrow: Lambda or Kappa, Lake or Warehouse or Lakehouse.

Why they look like alternatives

Presented side by side, each with a one-line definition, these patterns read as a menu. Ask instead what question each one answers, and they sort themselves into groups.

A sketched tree: Data architecture branching into where data lives, how data is processed, and who owns and governs it, with the relevant patterns under each branch.
Three independent decisions, not one menu.

Inside a branch the options really do compete: a table is a Warehouse table or a Lakehouse table, not both. Across branches they don't compete at all. "Lakehouse versus Data Mesh" is a question about two different things, like asking whether you would rather have a warehouse or a supply chain.

Medallion sits slightly outside this. It is a convention for organising layers inside whichever store you chose, and CDC is borrowed from application architecture and shows up at the edge of the platform.

How data is processed: Lambda and Kappa

Lambda

Lambda runs two processing paths over the same data: a batch layer that is slow and correct, and a speed layer that is fast and approximate. A serving layer merges them.

A sketched diagram: a data source splitting into a batch layer and a streaming layer, both feeding a serving layer and then analytics.
Two paths over the same data, merged at the serving layer.

It solves a real problem: you get historical accuracy and low latency at the same time.

The cost is the thing to weigh. You maintain the same business logic twice, in two engines, with two sets of bugs. Anyone who has reconciled a batch total against a streaming total at three in the morning knows what this costs, and it is the reason Kappa exists.

Kappa

Jay Kreps' argument in Questioning the Lambda Architecture was not that streaming is better. It was that keeping two implementations of the same logic is a maintenance burden you can avoid, because a durable log lets you reprocess history through the streaming path.

A sketched diagram: data sources feeding a retained, replayable event log, then stream processing, which serves both a historical store and real-time analytics, with a replay arrow curving back to the log.
One path, and history replayed through it.

Recomputation becomes replay: start a second job from the beginning of the log, let it catch up, swap over.

Kappa's own cost is less discussed. It needs a log you can retain and replay at the volumes involved, and the streaming engine has to handle workloads that batch engines find easy — large joins, wide aggregations, heavy backfills. Kappa is a good fit when most of your processing is genuinely incremental, and an awkward one when it isn't.

Choosing between them: if your workload is mostly incremental and your log retention is affordable, Kappa removes a whole class of duplication. If you have heavy analytical recomputation, Lambda's batch layer is doing real work that streaming would do badly.

Where data lives: Lake, Warehouse, Lakehouse

The warehouse gave structure, SQL and transactions but insisted on schema before you could store anything. The lake took anything at all but gave up transactions, performance and, often, order. The Lakehouse papers argued you could have both: open file formats on cheap object storage, with a metadata and transaction layer on top.

A sketched diagram: batch and streaming sources landing in open storage, refined through Bronze, Silver and Gold tiers, then served to BI, ML and applications.
One copy of the data, refined in place and read by every engine.

Microsoft Fabric's OneLake and the Databricks lakehouse are both implementations of this idea. What to check when someone claims it: whether every engine reads one copy of the data, or whether the platform is still moving copies around behind a single user interface. That distinction is the whole argument, and it is what separates a unified platform from a portfolio.

Medallion is a layering convention, not an architecture

Bronze for raw as-ingested data, Silver for cleaned and conformed, Gold for business-ready aggregates. It is a naming convention for a refinement pipeline, and it works equally well in a warehouse, a lake or a lakehouse.

It is worth saying plainly, because Medallion is often set beside Lambda and Mesh as though it were the same kind of decision. It isn't. You can adopt it on a Tuesday afternoon; the others reshape teams and budgets.

How data is owned: Mesh and Fabric

Data Mesh

Data Mesh is an operating model that happens to have architectural consequences. Zhamak Dehghani's principles are domain ownership, data as a product, a self-service platform, and federated computational governance.

A sketched diagram: a self-service platform above sales, finance and supply chain domain teams, each publishing a data product consumed by the wider organisation.
Domains own their data and publish it as a product.

The honest caveat: Mesh solves an organisational bottleneck, where one central team cannot keep up with demand from many domains. If you don't have that bottleneck, Mesh adds coordination cost and gives nothing back. It is the pattern most often adopted for the wrong reason.

Data Fabric

Data Fabric attacks a different problem: data spread across systems you are not going to consolidate, and possibly cannot. Instead of moving it, you put a metadata, governance and integration layer across it.

A sketched diagram: on-premises, cloud, SaaS and lake systems under a data fabric layer providing metadata, governance and integration to data consumers.
A layer over systems you are not going to consolidate.

Mesh is about who owns data. Fabric is about how you reach it. An organisation can sensibly do both, and the two are frequently confused because vendors sell products named after each.

Change data capture, at the edge

CDC reads a database's transaction log and propagates changes downstream, so an analytical store tracks an operational one without a nightly full extract.

It belongs in this discussion, but not as a peer of the others. CDC is usually the first thing you build and rarely a strategic choice — it is how data gets from the systems that run the business into the platform that analyses it. Where it does become architectural is in what it feeds: point it at a replayable log and you have the foundation Kappa depends on.

The question each one answers

The question you are actually asking

The pattern

How do I process batch and real time together?

Lambda

How do I avoid maintaining two copies of the same logic?

Kappa

Where do I put large volumes of varied, raw data?

Data Lake

How do I get SQL, transactions and governance over that data?

Lakehouse

How do I organise refinement inside whatever store I chose?

Medallion

Who owns data when one central team cannot keep up?

Data Mesh

How do I govern data I am never going to consolidate?

Data Fabric

How do I keep an analytical store in step with an operational one?

CDC

What a real architecture looks like

Because they answer different questions, they stack. A common and coherent combination:

A domain team owns a data product. Operational changes arrive through CDC, land in a replayable log, are processed Kappa-style with no separate batch path, refined Bronze to Silver to Gold inside a Lakehouse, and published as a data product under a Data Mesh operating model.

Every pattern named there is doing different work, and none of them is in conflict with another. If a diagram of your platform can't be described that way — naming which question each layer answers — that is usually a sign it contains patterns adopted because they were fashionable rather than because something needed solving.

The reverse test is more useful still. For each pattern in your design, name the problem that would return if you removed it. Anything that survives that question is carrying its weight.

Platforms in this article

Compare their capabilities side by side →

Tags: architecture, lambda, kappa, data-mesh, data-fabric, lakehouse, medallion, cdc