The Artificer by Loopit
Architecture Essay

Objectville's central warehouse: one common contract the whole chain subscribes to

By Santiago Coca · 13 min read · Part 07 of 15
Dark industrial warehouse with symmetrical shelving and dim blue light
A versioned central warehouse: every domain signs on to the common contract and contributes its own piece.

The warehouse robot arm

Last week I signed off with a robot arm that served pizza with noodle sauce whenever the boxes in the warehouse were shelved wrong. The pizza came out with academic confidence — and badly made. The question I left hanging was this: how do you organize the warehouse so the arm can't get it wrong?

Today I'm answering it. And the best way to see it is from inside a franchise that already gets it right.

Monday inspections

Monday morning at any Objectville Pizza Store branch. The inspector walks in. Unannounced, as any decent inspection should be. He asks for three things: traceability for the weekend's raw ingredients, the oven temperature log, and the allergen list served with every dish.

The manager doesn't head to the back office to dig through binders. She doesn't call headquarters. She doesn't open a Word doc. She opens one system, opens two tabs, and shows him.

Fifteen minutes later the inspection is closed. No arguing. No finger-pointing. None of that deer-in-the-headlights look a small business gets when someone walks in carrying a clipboard.

The interesting part isn't that the inspection was passed. The interesting part is what the branch didn't have to decide to pass it. It didn't decide what counts as a batch of mozzarella. It didn't decide which code identifies the customer who ordered the Chicago Style. It didn't decide whether Monday's batch MZ-7842-IT is the same one that arrived Tuesday as MZ-7842-IT-2, or whether Wednesday's chopped onion counts as “fresh” or “prepped”.

None of that is the branch's call. The franchise has already decided it. And it has decided it the same way for all forty branches — not case by case, not negotiated manager by manager, not according to neighborhood habit. One manual. One common contract. Forty managers with autonomy over everything the contract doesn't pin down.

Last week I asked you how you implement a data pantry. Today I'm answering that, and the robot arm, at the same time. The first thing you'll notice when you walk into the warehouse is that it doesn't distribute one thing, it distributes two — and that nuance, small as it looks, is what separates a franchise that scales from a chain that falls apart at the first crisis.

Axis 1 — Headquarters distributes two things, not one

In most chains, headquarters distributes ingredients. Flour, tomato, mozzarella, oil. It audits the quality of every batch, negotiates prices with the supplier, and at the end of the day sends out a truck full of labeled boxes. Each branch gets its boxes, identifies them as best it can, and fends for itself fitting them into its own process.

Box of mozzarella with a label signed by headquarters: batch, supplier, allergens and validation rule
Headquarters signs the label. The branch applies it — without renegotiating identity.

That's already a lot better than every branch ordering its own flour. But it's not enough.

At Objectville, headquarters does something more. Along with the audited ingredient, it also distributes its identity label. The branch receives the mozzarella and, hanging from it, a tag that reads: “Mozzarella di bufala — batch #2026-04-23-1742 — certified supplier — allergens: dairy — good until Friday — franchise identifier: MOZ-7842”. That tag is signed by headquarters, not by the supplier. The branch doesn't negotiate it. It applies it.

Why two things and not one? Because when a pizza comes through the door at three different branches — Rome, Boston, Chicago — and the inspector asks “which batch did that ingredient come from?”, all three answer with the same identifier. All three can trace back to the same supplier. All three can, if there's a recall for contamination, pinpoint exactly which pizzas used which ingredient and notify their customers.

Identity arrives signed by headquarters. The branch doesn't negotiate it. It applies it.

Carry that over to data and it works exactly the same way. The data franchise doesn't just distribute “ingredients” — raw table schemas, warehouse access, copies of the source CSVs. That's what an immature architecture distributes — “here's BigQuery, knock yourself out” — and it's exactly what makes every domain reinvent its own customer, its own product, its own idea of what counts as a “contract”.

The data franchise distributes two things:

  1. Audited ingredients — the what: a shared pantry with the canonical ingredients of the whole chain. Customer. Product. Contract. Point of sale. Each with a common code, one single way to identify it.
  2. The signed label — the ingredient's full trail: every time a domain loads data into the pantry, the platform signs not only who the ingredient is, but also which order it came from, when it arrived, what characteristics it has and which validations it has passed. The customer with tax ID NIF/CIF that Policies loads and the customer that Claims loads, if they're the same customer, carry the same signed label. Not by coincidence. By contract.

That signed trail is what makes it possible, when an auditor asks “which customers have open policies and open claims at the same time, and since when?”, for the answer to exist — and to be provable. No coordination meetings. No central team reviewing. The two branches speak the same language because the warehouse signed off on the translation the moment the data came in.

Axis 2 — Hub, satellite and link (and, finally, the full name)

What we just saw — the canonical catalog on one side, the signed label with the full trail on the other — has had a name of its own for two decades. The catalog is identity: what counts as a customer across the whole chain, and where that definition lives. The signed label is state: what we know about that customer, what score they have, what address they live at, what phone number they had last year. The two kept apart.

That separation has had a first and last name for two decades. Until today I had called it by its function — because I wanted you to understand why the separation is right before the name got in the way. Today I'll give it its full name.

It's Data Vault 2.0.

The identity of each shared entity — the customer, the product, the contract — lives in a piece called a hub. The hub is deliberately austere: it holds the business key, a technical key derived through a deterministic hash, a stamp of when it was first seen and which source it came from. Nothing else. That austerity is on purpose: the hub defines what counts as a customer, not what we know about them.

State lives in a piece called a satellite. The satellite hangs off the hub and describes versioned attributes. If the customer changes address, the old address isn't overwritten — a new row is inserted into the satellite with a timestamp, and the old one stays alive in its own era. If Policies stores customer attributes — onboarding date, score, commercial segment — they live in its satellite. If Claims stores others — claim report history, fraud indicator, adjusted premium — they live in another satellite, hanging off the same hub.

And here's the property no traditional warehouse solves:

Policies and Claims can add, evolve and refine their satellites in parallel, with no human coordination between teams, without stepping on each other's toes, without breaking the branch next door. Each satellite is an independent slice of the model that plugs into the shared hub. Concurrency is structural — it's not something you solve with a coordination committee; it's something that's solved because the model is designed so two domains don't need to meet in order to add their part.

That property — parallelism in both work and loading: two domains working at the same time, without coordination, without stepping on each other — is the first thing any serious Data Vault 2.0 implementation has industrialized for years. It's the methodology's common vocabulary, nobody's patent. What matters today isn't the machinery that automates it — what matters is why the hub/satellite split is what makes the chain possible.

And while all this is going on, another property is quietly at work: nothing gets deleted. The Policies satellite holds today's address, but also last year's, and the year before's, each with its own validity range. That's called insert-only — you only insert, you never overwrite — and it's what makes the chain's memory hold up effortlessly. When, six months from now, someone discovers that a record loaded on March 14 was wrong, there's no need to rebuild anything from old paperwork. The incorrect row is still there, its timestamp closed by a later correction.

Another consequential property follows from that, also too dense to open up today: reprocessing = replay. You can reproduce the March 14 calculation with the rule in force that day precisely because the versions were never destroyed. Next week I'll open it up properly. For today it's enough to know that it lives in the model, and it's a consequence of insert-only — not a module bolted on top.

The hub holds identity. The satellite holds state. And there's a third piece missing, the one that connects entities to each other: the link.

When a record arrives at the warehouse saying “customer Pepe holds policy P-1234”, that's neither an attribute of Pepe nor an attribute of the policy. It's a relationship. And relationships, in this methodology, aren't stored inside the entities — they're stored in a piece of their own. A Customer-Policy link is the pair of identities + a timestamp of when that relationship was seen + the source that brought it. Nothing else.

Why do you need a whole piece just for that? Because links emerge automatically from parallel loading. Policies loads its Customer and its Policy. Claims loads its Customer and its Claim Report. Nobody is talking to anybody. But the Customer-Policy link that Policies loads and the Customer-Claim Report link that Claims loads both connect to the same Customer hub — without anyone having to say so, without a coordination meeting.

That magic rests on two operational properties worth naming before we go on, because without them there's no parallelism at all:

Business keys. These are the business keys that give identity to each canonical hub — the customer's tax ID, the policy code, the book's ISBN, the product's SKU. The franchise decides which business key gives identity to each hub, and that decision lives in the Shared Kernel — it's not up to each domain to negotiate. “In this chain, a customer is identified by NIF/CIF — period.” When two branches talk about the same NIF, they're talking about the same customer. Nothing more is needed.

Deterministic hashing. The same NIF, normalized, always produces the same technical key — no matter which branch loads it, when, or from which system. That's why when Policies and Claims load the same NIF, the technical key they produce is identical — and that's why their links hook onto the same hub with no human coordination.

The model's three structural pieces wrap up here, along with their owners:

Piece What it's for Who decides
Hub Identity — who's who in the chain Headquarters (Shared Kernel)
Satellite Contextual state — versioned attributes Each domain (autonomy)
Link Relationships between entities Emerges mechanically from parallelism

Together, the three are what let forty domains load at the same time, against the same hubs, with no meetings, without stepping on each other, without the branch next door breaking.

But all of this — hub, satellite, link, business keys, hashing, parallelism, insert-only — rests on something deeper. It rests on a common agreement that the franchise signed before the onion ever left the field. And that common agreement has a technical name that isn't ours — it's borrowed from software design, like so many other pieces of this series.

Axis 3 — Shared Kernel: the common contract nobody negotiates

Back to the Objectville branch. The manager decides which pizzas go on her menu — the classic Neapolitan, the Chicago Style, the California Dreamin' in LA, a local pizza with payoyo cheese in Cádiz. She decides which local suppliers complement the shipment from headquarters. She decides her opening hours. She decides which customers she wants to win over and how.

What she does not decide:

  • What counts as a batch of mozzarella di bufala.
  • Which code identifies a regular customer.
  • Which identity label each ingredient she receives carries.
  • How traceability chains together when the same customer buys at two different branches.

Headquarters decides that. And it decides it once, for the whole chain. That small piece, deliberately minimal, is what Domain-Driven Design (Eric Evans, 2003) gives a technical name: the Shared Kernel. A minimal common contract that two or more domains share without negotiating it.

The data franchise works because the platform establishes an executable Shared Kernel: the chain's canonical entities (Customer, Product, Contract), the warehouse's three unbreakable rules (separate identity from attributes, never overwrite, identify deterministically), and the minimum validation rules a record has to meet to get in. Everything else — which satellites to add, which local rules to apply, which products to compose — each domain decides on top of that kernel.

What sets this apart from any traditional governance program is one word: executable. The Shared Kernel isn't a PDF pinned to SharePoint that “we all ought to respect”. It's code the platform runs every time a record comes in. The rule that says “a Spanish customer's NIF has to pass the official validation algorithm” doesn't live in a best-practices guide someone will review someday: it lives attached to the Customer hub and runs on every load. If a record comes in that doesn't meet it, it doesn't get in. If a new variant shows up — an identifier from a new legal regime — the system detects it and forces the rule to be refined. (I'll come back to this in Axis 4.)

That — an executable Shared Kernel — is exactly what Zhamak Dehghani identified as the fourth pillar of Data Mesh: federated computational governance, turned into code.

Let's recall the four pillars she described years ago: domain ownership of data, data as a product, a self-serve platform, and federated computational governance. The most operational question the model has had to answer since day one lands squarely on that fourth pillar: “how do you federate governance without every domain reinventing the rules? How do you compute governance without the platform turning into a central bottleneck?” The concrete answer — the one you can implement next Monday — is the one we're looking at: an executable Shared Kernel that the platform guarantees by construction and that domains extend with their satellites.

This is also where a question several of you left in the comments comes in, about canonical identity and contextual projections: is there one customer, or are there many? In the model, the hub/satellite split settles it: there is one customer — the one in the canonical hub, identified by the entire chain with the same key. And there are many projections of that customer — the one Policies sees, the one Claims sees, the one Marketing sees — each in its own satellite. Identity is one. Description is contextual. That symmetry is what keeps the chain from breaking into silos.

And this is where today's circle closes: Data Vault 2.0 is, by construction, that executable Shared Kernel. Shared identity lives in the hubs. Contextual descriptions live in the satellites. Parallel loading comes built in from the factory. Traceability is structural, not an added layer. When you ask a methodology for exactly those three things — shared identity, parallelism, memory that doesn't get erased — you've described Data Vault 2.0 without knowing it. And the fact that the kernel is executable — living code, not a PDF — is what closes the circle.

“The federated computational governance Zhamak described has a technical name: Data Vault 2.0. The franchise, as architecture.”

Axis 4 — Living documentation: the rule that can't grow old in silence

We've reached the most misunderstood piece. A single sentence defines it: the rule is code, the code runs every day against reality, and when reality changes, the system detects it before anyone goes looking.

Screen with an alert for rows outside the catalog and a whiteboard extending the catalog to COOPERATIVA
The rule learns when the analyst validates the legitimate value; the PDF grows old in silence.

Go back to the Objectville warehouse. The Customer hub has a rule attached to it: “The known catalog for the tipo_cliente field lists the values PARTICULAR, EMPRESA and AUTONOMO. Any value outside the catalog is inserted with a review flag.” That rule isn't a note in a guide. It's code. It runs every time a record comes in.

One ordinary Tuesday, two rows arrive with tipo_cliente = 'COOPERATIVA' — a value outside the catalog. The rule flags them. An operator sees the alert. She calls the source domain. “It's legit — the advisory firm just started handling co-ops.” The rule gets extended, versioned, deployed. The flag is lifted. The chain never broke.

That's the seed. The difference between this alert and the PowerPoint that grows old in silence for two years until an auditor shows up — and everything the platform does so you can't ignore it — is what I'll land next week.

What's next

What happens when bad data arrives — truly bad data? Not the legitimate case from Axis 4 — the new regime the advisory firm just took on. That's evolution; you handle it by refining the rule. I'm talking about the supplier who changes its schema without warning and sends garbage. The mandatory field that arrives empty. The business key that shows up duplicated because somebody dropped the ball. Where does that data go? How do you reject it without losing the trace that it ever arrived? Who decides what counts as good data when there are forty branches and one single common contract?

The answer isn't bolted on top of the model — it's a property of the model. Today I planted it in passing when I talked about rules that run on every load. Next week I'll open it up properly.

I'll leave you the twist with a concrete image: a supplier truck arrives one Tuesday with five thousand liters of crushed tomato nobody ordered, mixed in with the good goods from the order. Batch MZ-7842-IT. Supplier BLUE-LINE. Do you send it back? Do you put it in quarantine with a visible label? Do you accept it marked “rejected” so that six months from now you can reconstruct what happened? There isn't one answer. There are three different severities, a doctrine signed two decades ago, a concrete piece of the model that holds the trail, and a chef who tastes the dish every day.

See you next week.