The Artificer by Loopit
Architecture Essay

Mother Sauces — how a platform grows without breaking what already works

By Santiago Coca · 11 min read · Part 12 of 15
Pots of mother sauces simmering under dramatic light in a dark kitchen
Mother sauces are certified shared calculations: each domain builds on them without breaking what already exists.

“A franchise that scales distributes mother sauces. One that doesn’t scale distributes recipes or raw ingredients. The balance isn’t an option — it’s the only workable position.”

Monday, 7:42 a.m.

It’s Monday morning at an Objectville Pizza Store branch opening for the very first time. The operator — Ana, hired three weeks ago, until then a waitress at the place on the corner — arrives at seven with a coffee and a notebook. She looks at the empty kitchen. She looks at the cold oven. And she switches it on.

By 7:42, the oven is up to temperature. The San Marzano tomato sauce is thawed in its pre-measured container. The graded cheese topping is waiting in its labeled tub. The product freshness rule, calculated early this morning with the formula validated by the whole chain, is taped to the front of the oven like a little paper sign: “optimal freshness between 5 and 7 minutes at 320°C — alert if batch exceeds 48 hours since preparation”.

At 7:45, Ana pulls the branch’s first pizza out of the oven: a classic Neapolitan with San Marzano sauce on top, mozzarella di bufala, basil leaves, a drizzle of olive oil. The same pizza they serve in Chicago, in Rome, in Cádiz. The same pizza — because the sauce is the same, the cheese is the same, the freshness rule is the same. The recognizable chain.

What matters about that first pizza isn’t the pizza. It’s what Ana didn’t have to do to get it out.

She didn’t calibrate the oven from scratch. She didn’t define the cheese supplier’s contract. She didn’t learn to make San Marzano sauce starting from the tomato. She didn’t hire a specialist who understands the chain’s common model. She didn’t haggle with central purchasing over which version of the mozzarella she gets. She didn’t write her own quality rules. She didn’t wire anything into the central machinery by hand. Everything was already done when she got there. The oven turned on, the sauce was pre-measured, the freshness rule was taped to the front. Her job on day one was to serve pizza, not to build the factory.

At a quarter to eight, the first customer walks in. An older gentleman, retired, who knows Objectville’s pizzas because he usually stops by the location on the big avenue when he visits his daughter. He orders a Neapolitan. He eats it. He pays. On his way out, he tells Ana something that says it all:

— This one tastes just like the one near my daughter’s. Good to have it here too.

That’s the recognizable chain. The pizza the customer recognizes instantly — the dough, the temperature, the cut, the product’s signature. And at the same time, it’s this branch. It’s here, not at the other location. Ana is Ana, not the operator from the other place. The branch has its own light, its own neighborhood, its own customers who will come in and build a story of their own. Unmistakable without losing its soul. Both at once.

What just happened is the heart of the third pillar of Data Mesh — the self-serve platform. A branch has opened without having to rebuild the factory. From minute one, it has consumed what the whole chain had already built. And it was able to start serving from the first customer, instead of spending six months setting up a project.

How is that possible? That’s today’s question. And it’s going to take me through an engineering principle that’s nearly forty years old, a calibration that principle needs when you apply it to high-volume data, and the operational answer the ecosystem has been using for two decades to solve it. How each branch builds its local pizza on top of these mother sauces — Strategy, Composition over Inheritance, the two dangerous extremes to avoid — I’ll tell you next Tuesday. Today we get to know the sauces.

Axis 1 — The Open/Closed Principle: the property that keeps the chain from breaking as it grows

There’s a software design principle that Bertrand Meyer formulated in Object-Oriented Software Construction (1988), and that anyone who has set foot in a serious software architecture knows by heart. It goes like this:

Open for extension, closed for modification.

In plain English: when a new need comes along — a new type of calculation, a new local rule, a new product variant, a new domain joining the chain —, the system should be able to absorb it by adding new pieces, not by touching pieces that were already working. The new piece gets designed, tested, published. The old piece stays where it was, same as before, without a single bit changed. What already worked keeps working. What’s new adds value.

Applied to a franchise, Open/Closed says: the common manual — the tomato sauce recipe, the mozzarella di bufala contract, the loyalty customer code — doesn’t get touched every time a new branch opens. Ana’s branch opens, consumes the manual, serves. The chain didn’t budge to take it in. The other 39 branches didn’t even notice. And still, the branch was able to start operating.

Carried over to the data model we’ve been building, Open/Closed materializes as a very concrete structural property of the machinery:

  • The catalog of canonical hubs is closed for modification. Once the franchise has decided that a customer is identified by tax ID (NIF/CIF), that hub doesn’t get touched — because 40 domains touch it in parallel, and breaking it would break the whole chain.
  • But it’s open for extension. If a new domain joins the chain — Marketing adds loyalty data, Risk adds credit scoring, Digital adds the customer’s web behavior —, a new satellite is hung off the Customer hub that was already there. Without touching anyone else’s satellites. Without touching the hub. Without coordination meetings.

This matters because it’s the most valuable operational property of the model. When a new domain needs to get into the warehouse, it gets in by hanging a satellite off the hub that was already there. The new satellite is new code. It has its tests, its quality rules, its load contract. But it doesn’t break anything that was already there. The other 39 loads that were already working keep working exactly the same: they read the hubs the same way, generate their satellites the same way, pass their quality rules the same way. The chain has grown without noticing.

That’s the property. 40 domains can each be adding things to the model on their own, with no meetings, no coordination, no stepping on each other’s toes. That’s Open/Closed applied to data. It’s what makes a workable franchise possible: a chain that grows without the new piece having to ask the old piece for permission.

Now comes the calibration. The principle works — and it works better the more disciplined the extension is. That deserves an axis of its own, because it’s where a lot of people coming from pure software into the data world lose their footing in the first six months of production.

Axis 2 — The sweet spot: Open/Closed in software vs. in high-volume data

Open/Closed is a solid principle. But extending in pure software and extending over high-volume data tables aren’t exactly the same operation. And understanding why is the difference between a data franchise that scales with its head on straight and a franchise paralyzed by its own micro-segmentation.

Comparison: a Customer hub with dozens of tangled joins versus six satellites grouped with discipline
Open/Closed with discipline: group satellites, don’t multiply joins unchecked.

In pure software, when you apply Open/Closed, the cost of adding a new class is basically zero at runtime. The new class gets compiled, loaded into memory, invoked when needed. The old class stays where it was. The system grows without paying a read cost on every operation.

In data, there’s an operational nuance. Every time you hang a new satellite off the Customer hub, that satellite is a potential JOIN whenever someone wants to rebuild the customer view. Every time you add a new quality rule, it’s one more operation on every load. Every time you split a domain into finer granularities — one satellite per team, one satellite per project, one satellite per whim —, you’re adding one more layer of read-time stitching whenever someone wants to query the whole entity.

If you extend without discipline — a new satellite every time a domain wants to add three attributes, one more subdivision every time a team changes scope —, you end up with a model where rebuilding a customer’s consolidated position means joining 47 tables, each with its bitemporal history, each with its versioning, each with its validity range. Open/Closed survives as a principle, but reading becomes a pain: adding is free, but reading costs. And the micro-segmentation that seemed to respect local autonomy starts getting in the way of the queries the business actually runs.

This isn’t theory. It’s what happens every week in real Data Vault projects set up without discipline. “If hub-and-satellite is good, more satellites must be better.” And two years later there’s a three-thousand-table model where every query touches fifteen and the analysts go back to asking for a CSV export “because querying the model is slow.” Governance has held. Performance has taken a hit.

The sweet spot exists, and it’s a matter of discipline. You don’t make it up. And the discipline comes down to two operational modeling rules — the ones that live on the same plane as Open/Closed — that separate a living franchise from a model paralyzed by its own granularity.

Modeling rule A — Group satellites by rate of change, not by the domain that contributes them.

Customer attributes that change every day — geolocation, credit score, commercial segment — live together in one satellite. Attributes that change once a year — date of birth, tax identifier, legal status — live in another. Attributes that change once in a lifetime — original sign-up date, first product purchased — live in a third. Mixing attributes with different rates of change in the same satellite forces you to read the full history of the slow ones every time you query the fast ones. The whole system foots the performance bill, on every query, every day.

This rule isn’t ours — it’s common vocabulary in any serious Data Vault book. Linstedt stresses it a lot. The reason it’s hard to apply is that it runs head-on into the intuition of “one satellite per team that contributes data”, which is the easy way to organize things but the worst in terms of read cost.

Modeling rule B — Limit the depth of read-time stitching.

If rebuilding the view of an entity — the customer, the policy, the contract — requires joining N satellites at runtime every time an analyst queries, there’s a threshold at which the model starts asking for an intermediate layer that crystallizes the result of the shared calculation the whole chain reuses. That intermediate layer has a technical name that will come up in the next axis, and it’s the exact operational answer to this problem.

How big is N? It depends on the engine (BigQuery can take more than Snowflake, Snowflake more than an on-prem Postgres) and on volume. But there’s a reasonable rule of thumb: if the business’s most frequent query has to touch more than 5–7 satellites of the same hub to rebuild a calculation the whole chain needs the same way, materialize. Don’t wait for it to blow up.

So far, two modeling rules. And now comes an important distinction, because there’s a third dimension that any senior engineer who has run Data Vault in production knows by heart — but that isn’t modeling.

🔑 What’s coming next week — and why it isn’t modeling

There’s a whole family of pieces in the Data Vault 2.0 catalog that exists to speed up serving the data when read costs start to pinch. PIT tables (Point-in-Time) that precompute as-of states so historical queries don’t recalculate validity ranges on every call. Bridge tables that precompute frequent paths between hubs across multiple links to avoid deep JOINs. Selective materializations of views the business consumes daily. All of this has existed for decades in the Linstedt catalog, and the whole serious ecosystem (Scalefree, AutomateDV, VaultSpeed, dbt vault) industrializes it.

But — and this is the important part — these pieces carry no business logic. They’re pure accelerators: they precompute paths or states so the query is cheap. They don’t decide anything the business should have decided. They’re not modeling, they’re physical plating: how what’s already modeled gets served to the consumer.

That’s why this whole family doesn’t come in today or next week. The Open/Closed doctrine wraps up today with two modeling rules. The pure-performance family — PIT, Bridge, materializations — I’ll cover later on, when we get to plating. Once you see that the dish goes out to the pass rather than being assembled in the kitchen, that whole family stops being a mystery.

And that connects to the modeling answer the DV 2.0 ecosystem has been using for twenty years. It has a name, it has a layer, and it has an operational unit we’re about to look at.

Axis 3 — Mother sauces: Vault Components and the Business Vault (the layer that carries logic)

Until now — through W7, W8a, W8b and W9 — we’ve talked about a single layer of the warehouse: the Raw Vault. The raw layer. The one that reflects what arrived from the source, untransformed. Canonical hubs, satellites with their versioned attributes, links connecting entities. Insert-only, total fidelity to the raw data. That layer is what makes traceability possible — “reproduce the March 14 calculation with the rule in force that day” — because it keeps the data exactly as it arrived.

But if all you have is a Raw Vault, every time the business wants to ask “how long has this customer been with us?” or “how fresh is this product?” or “what’s the consolidated score?”, someone has to recompute the formula from the raw data. In the pure software world, that would be trivial. In the high-volume data world, multiplied by thousands of queries a day and hundreds of different formulas, it’s an operational cost the system ends up paying in slowness or in duplicated calculations that diverge across domains.

The operational answer that Dan Linstedt introduced into the Data Vault 2.0 catalog quite a while ago, and that the whole serious ecosystem has been applying for years, is a new layer that lives on top of the Raw Vault: the Business Vault.

The Business Vault isn’t a recent invention, and it isn’t ours. It’s canonical vocabulary in the DV 2.0 ecosystem — Linstedt and the Scalefree folks describe it, AutomateDV and VaultSpeed and dbt vault industrialize it, it’s in any serious book on the framework. What I’m going to tell you is how it fits into the franchise metaphor and why it’s the exact answer to the previous axis’s sweet spot — on the modeling plane.

The Business Vault holds shared calculations the whole chain needs to use the same way, and that carry business logic inside. “Customer tenure”: calculated as the difference between first transaction and today, normalized to years — headquarters validates the formula, all 40 branches consume it the same way. “Product freshness”: calculated with the formula the food quality team validated — it factors in transport temperature, days since preparation, source batch. “Consolidated customer score”: a composite of three metrics (loyalty, credit risk, behavioral scoring) with weights validated by management.

Those calculations live exactly once in the Business Vault. Versioned — every change carries its timestamp. Tested — they pass automated tests before being published. Certified — someone accountable signs off on them. And every branch consumes them without rewriting them, injecting them into its local queries as if they were one more ingredient from its pantry.

Those are the mother sauces. Ready to combine.

The operational unit of the Business Vault is what the ecosystem calls a Vault Component — an encapsulated piece with a very precise contract. It carries its input contract (which hubs and satellites it needs), its output contract (which columns and semantics it produces), the guarantees it meets, its version, and its lineage. (The book chapter has the full detail, with the five elements spelled out, concrete contract examples and the mechanics of versioned deployment.)

When Ana’s branch opens and starts serving its first Neapolitan, it declares which Vault Components it will use (product freshness, customer score, customer tenure) and combines them with the chain’s common recipe. The branch doesn’t recompute freshness — it consumes it. It doesn’t recompute the score — it consumes it. It doesn’t rewrite the tenure rule — it injects it as is.

The sharp boundary — Vault Component vs. PIT/Bridge

Here’s where I have to say something that separates a serious doctrine from a muddled one. A mother sauce — a Vault Component — and a performance piece — a PIT table or a Bridge table — aren’t the same thing, even though both physically live as materialized tables and even though both are canonical in Data Vault 2.0. The difference is operational and sharp:

The difference is operational and sharp. A Vault Component (Business Vault) carries business logic: it decides the freshness formula, the score weights, the tenure rule — and that’s why it’s modeling, and it lives here, today. A PIT table (Point-in-Time) decides nothing: it precomputes as-of states for a hub so historical queries don’t recalculate validity ranges. A Bridge table doesn’t decide anything either: it precomputes the path between hubs across multiple links. And neither does a selective materialization: it crystallizes a frequent view to make the query cheaper. All three are plating — pure performance, no logic — and they come in later.

The line is this: if the piece decides something about the business, it lives in the Business Vault. If it only speeds things up, it lives in the serving layer. This boundary isn’t Linstedt’s — in the standard DV 2.0 taxonomy, PITs and Bridges are usually grouped inside the Business Vault. I draw it sharper on purpose: what carries logic is modeling; what only accelerates is plating. That line is what keeps governance clean.

Why does this distinction matter? Because when an organization confuses the two layers — and many do — it starts putting business logic inside PIT tables (“since I’m materializing anyway, I’ll throw this calculation in while I’m at it”), and that is exactly the recipe for a model that becomes ungovernable. A PIT table with logic is a Vault Component in disguise, without the certification, without the explicit version, without the lineage. The audit can’t find it. The change sneaks in. The chain starts to diverge without anyone noticing.

The doctrinal rule: any piece that makes business decisions lives in the Business Vault with its full Vault Component contract. Any piece that only accelerates lives in the serving layer — no logic, transparent, rebuildable from the Raw Vault and the Business Vault with no loss of information. That boundary is what lets governance hold even when the physical machinery has dozens of accelerators.

The Business Vault has its own sweet spot

Mother sauces solve the concrete problem the previous axis left open: when rebuilding a shared calculation means joining too many satellites at runtime, the Vault Component crystallizes it exactly once with its certified logic, and the whole chain consumes it instead of recomputing it. That’s exactly the intermediate layer I promised in Rule B of the sweet spot.

But watch out, because this layer has its own discipline — and it’s the same Open/Closed logic, one floor up. If you give in to the temptation of turning every calculation into its own Vault Component, you reintroduce the problem you set out to solve: components share underlying satellites and calculations, and when you break everything into tiny functions, those shared reads explode — every micro-component reads and joins the same things all over again. You’ve moved the stitching up a level, not eliminated it.

That’s why Business Vault discipline comes down to two questions:

  • When do I create a new Vault Component? Only when the calculation is something the whole chain reuses the same way — tenure, freshness, the consolidated score. A calculation for a single branch, or a local variant, isn’t a mother sauce. The ticket into the Business Vault is reuse, not elegance: if only one kitchen is going to use it, it’s not a mother sauce, it’s that kitchen’s dish.
  • Does it have a single purpose? Each Vault Component does one thing. The moment a component starts having more than one purpose — it calculates freshness and, while it’s at it, adjusts the score —, that’s your sign you’ve mixed two sauces in the same jar. Split it. A component with two purposes can’t be versioned, certified or reused without dragging along what you didn’t want.

The Business Vault isn’t the pretty calculation some architect thought was elegant. It’s the disciplined materialization — with logic, with a signature, with a version — of what the chain needs to reuse the same way. That turns the Open/Closed of axis 1 into something workable at scale: shared-calculation extensions, instead of scattering across 40 branches that each recompute their own, crystallize exactly once in an intermediate layer with its contract — without falling into the component micro-segmentation that sneaks read costs back in through the back door.

The serving layer — PIT, Bridge, materializations — gets built on top of both (Raw + Business) when needed, but it doesn’t make decisions. It’s transparent. And we’ll get to it in due time.

What’s next

The mother sauces are in the kitchen. What’s missing is how a local branch combines them without breaking the kernel — Denver with its altitude, Rome with its own.

Next Tuesday: Strategy, composition over inheritance, and the four elements that separate a real platform from theater.

And a manager in Chicago with a box of honey gold she’s been mulling over for three weeks.

Until then.