The mechanics are in place. The doctrine is missing
Last time, we left the factory with a signed-off structural property: nothing gets deleted and nothing gets overwritten. Insert-Only + Audit-Only-Never-Delete (Dan Linstedt, canonical Data Vault 2.0 vocabulary). We also left the versioning pattern inside every satellite — HASH_DIFF + VALID_FROM/TO windows; in the batch example, the version number labels the order of that specific record — with MZ-2026-PRO-014 on the table: two rows, two coexisting truths, one precise time window that closes one and opens the other.
Those are the mechanics. But mechanics aren’t doctrine. A factory with every versioning property in place can still use them badly — and in fact nearly every chain I see in production does use them badly — the first day a serious audit shows up. Today I’m closing the two questions still open since that Friday at the coastal branch:
- What exactly does it mean to reprocess the March 14 calculation? Spoiler: there’s a right answer and a wrong one. The difference between them isn’t cosmetic — it’s structural. And the right answer is the one that closes the whole pillar.
- How do you rebuild, before Monday, the exact list of pizzas served with batch MZ-2026-PRO-014 across five branches over six weeks — no project, no log dump, no coordinating with the platform team? Spoiler: one query against the model. The factory doesn’t build the proof; the factory is the proof.
One thing at a time.
Axis 1 — Reprocessing means reproducing, not recalculating: the rule of the moment, not today’s rule
There’s a word that comes up all the time in data governance conversations and is almost always misunderstood. The word is reprocessing. And the confusion is expensive.
For a very large share of data teams, reprocessing means recalculating with the current rule. In other words: if a March metric was calculated wrong — because the rule at the time was wrong — reprocessing means running the March calculation again, now with the corrected rule, and overwriting the previous result. The old metric disappears, the new metric takes its place. Reprocessing as a synonym for cleaning up the past with the knowledge of the present.
That interpretation is what sinks chains the day a serious audit arrives. And the reason is structural, not a matter of good intentions.
A serious audit — the one from the Italian farm consortium that certifies designations of origin, from a financial regulator, from a health authority that monitors drug traceability — asks things like this:
“Rebuild the March 14 calculation for me with the rule in force that day, not the rule in force today. I want to see what the machinery was telling you that day — not what you know today with hindsight.”
If the data team reprocessed in the sense of recalculated with today’s rule, the answer is: it can’t be done. The March 14 metric according to the March 14 rule is gone — it was overwritten by the reprocessing. The audit won’t accept today’s version; today’s version is a guess about the past, not a reconstruction of it.
What the chain’s factory does — and what the model side of the Federated Computational Governance pillar makes mandatory — is reproduce, not recalculate. The difference is sharp and operational.
Reproducing means: running the March 14 calculation against the data in force on March 14 (the rows with VALID_FROM ≤ 14-mar and VALID_TO > 14-mar or open), with the rule that was live on March 14. The query filters by the model’s time window, the rule applied is the version of the rule that was active that day, and the result of the calculation is identical to the original result. Not equal if you’re lucky — structurally identical, because the input (the data) and the algorithm (the rule) are the ones that were, not today’s.
Recalculating, on the other hand, would mean filtering by the March 14 time window but applying today’s rule — or worse still, reading the data in its most recent version and applying today’s rule. That’s not reproducing: that’s rewriting the past with the knowledge of the present. Structurally valid for some analytical tasks; structurally unacceptable when what you need to prove is what the model said at the time.
Why is the factory able to reproduce and not just recalculate? Because of two properties we already nailed down when we wrapped up the versioning mechanics:
- Data for any past date is still alive in its correct version. Native versioning guarantees that the row in force on March 14 can be retrieved by filtering the time window — it’s not lost data, it’s a specific row with its
VALID_FROMand itsVALID_TO. - Quality and calculation rules are versioned too. The rules live in the machinery — not in code scattered across scripts — and like anything else that lives in the machinery (the central warehouse distributes structures and rules), they’re subject to the same Insert-Only + Audit-Only-Never-Delete property. The February 2 validation rule for Pecorino Romano DOP lives with its own validity window. When it’s updated — because information about the lab’s faulty certificates forces a change in the rule — the old rule doesn’t disappear. Its window closes, the new one opens. And when you reproduce the March 14 calculation, the rule applied is the one live on March 14, not the one live today.
The mirror image of the Mediterranean chef: the auditor who asks about the past
This property — reproduce, don’t recalculate — isn’t exclusive to the kitchen. It lands exactly the same in any domain where a rule evolves and an auditor asks you to account for the past. And it’s worth grounding with a case from the data world that’s been rolling through this series for a while now: the chef’s pass and a live rule in the Accounts domain: “If a cancelled contract has a balance of €500,000, I don’t buy it.” The rule was born simple — “if status = CANCELLED, then balance = 0” —, severity CRITICAL. It lived in the machinery. It ran every night. And it rejected — correctly — contracts where a source failure had left a nonzero balance after cancellation.
The catalog then did what it should when a domain learns. That same rule evolved. Following an operational approval to round off residual balances from accrued interest, version V2 went into production on April 15: “if CANCELLED then balance = 0 OR balance < 5”. V1 was closed with its VALID_TO on April 14 at midnight. V2 opened with its VALID_FROM on April 15 at 00:00:01. Zero sync gap — the catalog learned, the machinery consolidated.
A Friday in June comes around. An internal control auditor from the Banco de España asks for something specific:
“Rebuild which cancelled contracts with a nonzero balance were flagged as rejected on March 1, according to the rule in force on that date. I want the exact list — not today’s.”
Here’s the fine cut. The question has two possible answers, and only one is correct.
The wrong one — the one you get from recalculating with today’s rule (V2) — would return the list of cancelled contracts with a balance > €5 on March 1. That’s what the machinery would reject today if a batch arrived with that information. But it isn’t what the machinery rejected on March 1. That day rule V1 was live — balance = 0 —, so any cancelled contract with €1.73 or €4.27 was on the rejected list. Recalculating with V2 leaves them out. The answer under-reports the past by applying a rule that didn’t exist yet.
The right one — the one you get from reproducing with the rule of the moment — filters the data in force on March 1 (contracts with a VALID_TO after that date) and applies the version of the rule live on March 1 (V1, in force from Q3 2025 to April 14, 2026). It returns the exact list — including the €1.73 and the €4.27 — that the machinery rejected that day. The audit is closed in one query. The machinery knows which rule was live on which day because rule V1 wasn’t deleted: its VALID_TO was closed on April 14. Just like the certificate for the cheese batch.
And there’s the pattern. The same structural property works in the kitchen and in banking. The Mediterranean chef forced V2 of a rule — the chain learned going forward. The coastal manager asks to rebuild the past under the rule of that moment — the chain reproduces going backward. Two questions, two different branches, two different domains, one machinery. Both answers come from the same query — only the time filter and the version of the rule applied change.
This closes a callback that’s been open since this series entered the technical block, and it has a specific name. Federated Computational Governance — the Data Mesh pillar formulated by Zhamak Dehghani — proposes that governance isn’t a program of human committees approving exceptions against the current. It’s executable governance, carried by code that travels with the data and lives in the model. The factory’s memory is the structural face of that principle: governance doesn’t just decide what gets in and what stays out today (rules at the door, quarantine, a quality log when something smells off), it also keeps alive the record of what it decided yesterday and what it knew yesterday, so it can answer questions about yesterday without passing them off as questions about today.
In industrial terms, it’s the same property that separates the factory whose ovens remember what temperature they ran at last month from the factory whose ovens only know what temperature they’re running at now. Data Vault 2.0 as an architecture lets you reproduce by construction, with no extra work. And that’s the property that closes the Federated Computational Governance pillar.
Axis 2 — Tracing a spreading error: batch MZ-2026-PRO-014 and Monday’s query
Back to Friday. The coastal branch manager has hung up, headquarters has until Monday to rebuild the traceability of batch MZ-2026-PRO-014, and the concrete operational question is this:
At which branches, during what window, were pizzas served with batch MZ-2026-PRO-014 labeled as Pecorino Romano DOP — when it was actually pasture pecorino with no designation of origin?
Monday at nine in the morning, the answer arrives.
On Saturday afternoon someone at headquarters had fired off a single query against the model. On Monday nobody has to rebuild the incident by hand: the machinery already says where the batch slipped in and for how long it circulated under the wrong label. The map returns five branches; the coast takes the brunt. North Coast carries the weight: 142 pizzas between February 2 and March 18, 89 tables affected, 27 customers with named invoices — the longest, densest trail. South Bay shows the same pattern with less volume but the same end date: 98 pizzas from February 9 until the certificate correction, 64 tables, 19 named customers. West Central, East Station and Plaza Mayor round out the picture with shorter stretches; the model returns them in the same query.
Chain summary (all five branches together): 371 pizzas with MZ-2026-PRO-014 under the DOP label the lab certified wrongly; 253 tables affected; 70 customers with named invoices in the cut. Overall incident window: from the first service with the batch to the March 18 correction — that’s what’s going to matter when the Italian farm consortium asks for dates.
That answer isn’t hypothetical or a guess. It’s the result of that one query against the model — no project, no dump of old logs, no coordinating with the platform team or waiting for a maintenance window. The query crossed three satellites (the batch, the menu, the served pizza), two hubs (batch and branch) and one link (the one connecting batch to served pizza and branch). Plain SQL syntax over a model built with the Data Vault 2.0 methodology. The same query profile the team uses to answer “how many pizzas did we serve in March at North Coast?” — except this time the condition filters on a specific batch and on its validity window as DOP, before the March 18 correction.
And here the factory gains more than speed. It gains provability. Because when the Italian farm consortium asks for traceable documentation of the incident — “prove that between February 2 and March 18 the chain didn’t know the batch was non-DOP” — the answer doesn’t require building a post-mortem from outside the model. The answer lives in two rows of the batch satellite: version 1 of the record with a DOP certificate and VALID_FROM = 02-feb-2026 09:14:08, and version 2 of the record with a NON-DOP certificate and VALID_FROM = 18-mar-2026 11:42:33. Both signed by their operators, both untouched since insertion. The factory doesn’t build the proof — the factory is the proof.
And what about quality at the door — that log where the factory notes severity and context when something doesn’t fit, without mixing it into the operational chatter of the business satellite? This is where the other piece comes in. That piece — in Data Vault 2.0 it’s usually modeled as a quality satellite — logged an entry on March 18 with severity WARNING — “retroactive certification correction: batch MZ-2026-PRO-014, reason: documentation error at the supplier’s lab, requires notification to branches that consumed the batch in the 02-Feb to 18-Mar window”. That entry is what automatically triggers the branch-by-branch breakdown and the total of 371 pizzas. The quality log at the door and the model’s memory aren’t two systems emailing each other — they’re two faces of the same machinery, and when a retroactive correction lands in that log, the traceability comes out of the model with no extra work.
Is there anything else to do on Monday, besides notifying the branches and preparing the documentation for the consortium? Yes — and this is where the day’s foundational lesson lands.
You have to review the DOP certificate validation rule that passed the batch on February 2. The rule at the time was reasonable — it trusted the accredited lab’s certificate and the official visual seal. The reality is that trust wasn’t enough, because an accredited lab can issue a wrong certificate (in this case, it did). The new version of the rule — version 2 of the rule, not of the batch — goes live on March 19 and adds a step: “cross-check the certificate number against the central registry of the designation-of-origin issuing body before accepting the batch.” The old rule stays in the model with its VALID_TO closed on March 18. The new rule comes in with its VALID_FROM on March 19 and an open VALID_TO. The chain doesn’t break what it validated in February — that was valid under the rule of the moment. The chain learns and improves the rule going forward, without touching the past.
And the subtlest but perhaps most important property: improving the rule doesn’t force you to throw away anything already in the model. The satellite’s structure doesn’t change. The old data stays where it was. The rule improves — that comes in as a new version of the rule, not as a new data structure. The machinery is open to extension (rules can improve) and closed to destructive modification (the improvement doesn’t require touching what’s already built). We haven’t called that property by its canonical name yet — it comes from a specific author and will land formally in the next chapter of the modeling blueprint, anchored to that author’s last name — but the seed is planted here. If you’ve noticed that this way of evolving the chain is what separates a factory that lasts from the six-month project that gets redone every time a rule changes, you already have half the principle in your head.
There’s one image that closes the day. The whiteboard in the central kitchen, Monday at nine in the morning, after notifying the five branches and sending the traceability dossier to the Italian farm consortium. Under the breakdown the query dumped onto the board, someone has written in chalk a sentence an outside visitor would understand with no context:
“Nothing gets erased here. Everything gets versioned. That’s why the factory remembers — without having to take notes.”
It’s not marketing. It’s the factory’s structural property.
What’s next
The Federated Computational Governance pillar is closed here.
The question that decides whether the chain scales beyond five branches is a different one: when it opens a new location in a city where Objectville has never been — what does the branch do on day one? If the answer is setting up the oven, the contracts and the rules from scratch, every opening is a project. Projects don’t scale — they pile up.
The chain that scales does the opposite: oven already installed, sauce pre-measured, rules signed off, memory available from day one. A new operator shows up at seven with a coffee and a notebook. By 7:42, the oven is already hot.
See you next Tuesday.