Last week we opened the pantry. We saw that storing data permanently and in an organized way — separating identity from state — is what lets you answer an audit in minutes and reprocess without depending on anyone.
But there's something I didn't tell you.
Having the perfect pantry doesn't mean you can serve a single dish.
“We have all the data. We don't have a single answer.”
On a banking project, the team had done things right. The raw data was organized, labeled, with full history. The pantry was spotless. When the regulator asked for the figure as of March 15, they found it in seconds.
But then a different request came in: “We need the average daily balance of every account over the last quarter.”
Silence.
Not because they didn't have the data. They had all of it: every balance, every day, every account. The problem was something else. Balances didn't arrive every day. On some days, some accounts had no activity. Was the balance that day zero? No — it was the same as the day before. But in the data, that day simply didn't exist. It was a gap.
And an average balance is not an arithmetic mean. You can't add up the balances you have and divide by the number of records, because you'd be ignoring the days with no record. A customer who held a balance of €100,000 for 89 days and had a €1 movement on day 90 doesn't have an average balance of €50,000. Their average balance is ~€98,889. But if all you have is the two records (100,000 and 1), the arithmetic mean says 50,000.
The team had a senior engineer who knew the problem. He wrote a 400-line SQL script that generated a full calendar, filled the gaps with the last known value, and then computed the average weighted by the actual days. It worked. For one account.
But they had 200,000 accounts. With 15 indicators each. And the regulator needed it for 6 quarters.
The script took 4 hours to run. Nobody else on the team understood the logic. And when, three months later, the regulator changed the definition (“the average balance is now calculated on calendar days, not business days”), the engineer had to rewrite 200 of the 400 lines.
But that wasn't the worst part. The worst part was that, in the meantime, another team at the same company needed to calculate the average balance of a different kind of indicator. Since they didn't know a script already existed, they wrote their own. 350 lines. With slightly different logic. Different on purpose, or by mistake? Nobody knows. Now there were two versions of the “average balance” in production, with different results, and nobody could say which one was right.
The data was there. The answer wasn't.
Having the ingredients in the pantry doesn't mean you can serve a dish. You need to prep those ingredients. That's mise en place.
Mise en place: the invisible work that changes everything
In a professional restaurant, nobody peels onions during service. Nobody fillets a salmon when the server shouts “Table 7!” Nobody measures spices with the dish already on the flame.
All of that happens beforehand. That's the mise en place: the quiet prep that happens before the first guest walks in. Every vegetable washed and cut. Every sauce reduced and ready. Every protein portioned and labeled. Everything in its bowl, measured, prepped, good to go.
It's the least glamorous work in the kitchen. Nobody posts their mise en place on Instagram. But ask any chef: without mise en place, service is chaos. With mise en place, service flows. The difference between a kitchen that turns out 200 plates a night and one that collapses at table 15 is the prep nobody sees.
In data, mise en place is exactly that: the step where raw ingredients from the pantry are turned into information the business can use. Where the balance on day 15 becomes the average balance for the quarter. Where gross premium becomes net premium. Where rows of transactions become the portfolio assigned to each branch office.
It's the layer where business rules are applied. And it's the layer where, in most companies, everything breaks.
The anti-pattern: the chained mise en place
Let me describe how mise en place works in most of the data warehouses I've seen.
There's a pipeline. It starts by extracting data from the source system. That data lands in a temporary table. From there, a script cleans it and drops it into another table. Another script picks up that table, applies a business formula and produces a third table. Yet another script aggregates, pivots, and leaves the result in a view that feeds the report.
Each step depends on the previous one. Each table is the input to the next script. It's a linear chain, like a row of dominoes.
And it works. Until someone nudges the first domino.
One Friday, the source-system team adds a new field to the balances table. Nothing existing changes — they just add a column. “It's backward compatible,” they say. And technically, it is.
But the cleansing script has a SELECT * that now brings in one extra column. The temporary table changes schema. The calculation script, which expected exactly 47 columns, fails. The aggregation script doesn't even start because its input doesn't exist. Monday's report comes out empty.
One new field. Backward compatible. And the whole chain breaks.
The problem isn't the SELECT * (which, yes, is bad practice). The problem is coupling. Every link in the chain knows the internal details of the one before it. If the previous one changes — even in a “compatible” way — the next one can break. It's a house of cards dressed up as a pipeline.
And the worst part isn't that it breaks. The worst part is that you don't know what's going to break until it breaks. Because the dependencies are implicit. They aren't declared anywhere. They live in the SELECTs, the JOINs, the WHEREs of scripts nobody documents and only the person who wrote them understands.
I've seen teams spend more time fixing breakages in their mise en place than building anything new. Every Monday starts with the same question: “What broke this weekend?” Not because the data is bad. Because the transformation chain is fragile.
Remember the 40% of reactive time we mentioned in Week 1 (Monte Carlo Data figure, 2023)? A big chunk of that time goes into exactly this: figuring out what broke, where, why, and fixing it before someone opens a report and sees absurd numbers. That's not engineering. It's plumbing. And it's the direct consequence of a chained mise en place.
Your mise en place is a castle of dominoes. And every Monday is an earthquake.
Functions, not chains
Back to the kitchen. And think about what happens when a cook starts cooking straight away, with no prep.
In go the carrot, the tomato and the onion. Start the sauté. So far, so good. But suddenly a new order comes in: same dish, no tomato. What now? You can't pull the tomato back out of the pot. It's already mixed in. Already coupled. You have to start from scratch with another pot.
That's exactly what happens when you couple transformations at the start of the pipeline. If your first script mixes extraction, cleansing and calculation into a single chain, any variation forces you to start from scratch. Another team needs the same data but with a different calculation? They can't reuse your pot. They have to build their own from the ground up. And that's how the 10 versions of the average balance are born.
If instead you have the carrot in one bowl, the tomato in another and the onion in a third, you can combine them however you like. Dish with tomato: grab three bowls. Dish without tomato: grab two. The bowls don't change. The combinations do.
In a professional mise en place, the bowl of diced onion doesn't know the bowl of sauce exists. The bowl of sauce doesn't know the bowl of portioned meat exists. Each prep is independent. It has its input ingredient, its process and its result. If tomorrow the chef decides to switch the onion cut from brunoise to julienne, the sauce never finds out. Because the sauce doesn't depend on how the onion is cut. It depends on its own ingredients.
And if the chef wants to add a new bowl — some pickled vegetables that weren't on the menu before — she sets it on the counter next to the others. No need to rearrange the whole mise en place. No touching the bowls already there. She just adds one.
You can add new preps without modifying the existing ones. And each prep is an independent function: it takes defined inputs, produces a defined output, and has no side effects on the rest.
So why don't we do this with data?
Think about our average balance. What does it need as input? A time series of balances. What does it produce? An average balance per period. It doesn't need to know where the balances come from. It doesn't care whether they're from a bank or an insurer. It doesn't care whether the source has 10 columns or 200. All it needs is a date column, an amount column, and a key that identifies the account.
It's a function. Inputs → process → output. And if tomorrow you need the average balance of another indicator, you don't rewrite the 400-line script. You use the same function with different parameters.
Same with filling time gaps. What does it need? A time series with possible gaps. What does it produce? A complete series, no gaps, with a defined fill policy (repeat the last value, set to zero, interpolate). It doesn't care what the data represents. It's a function.
And they compose: first you fill the gaps, then you compute the average balance over the complete series. Two functions. Each one independent. Connected by a simple contract: “the output of the first is a complete time series, and the second expects exactly that.”
If tomorrow you change the fill policy from “repeat last value” to “linear interpolation,” the average-balance calculation never finds out. Because it doesn't depend on how the gaps were filled. Only on the series being complete.
And if the day after tomorrow you need a new calculation — say, maximum intraday variation — you add a third component that receives the same complete series. You don't touch the first two. You don't need to coordinate with anyone. You don't break anything. You just add a bowl to the counter.
Now think back to the problem at the start: two teams with two versions of the average balance. Remember how we couldn't tell whether the difference was intentional or a mistake?
With independent bowls, the answer is clear. If the second team needs an average balance calculated differently — say, with calendar days instead of business days, or with a different fill policy — they don't modify the existing component. They create a new one. With its own parameters. Its own contract. Its own name.
And the two coexist. No conflict. No ambiguity. The Treasury team keeps using SaldoMedio-DiasHabiles. The Regulatory Reporting team uses SaldoMedio-DiasNaturales. Each one takes its ingredients from the pantry and its preps from the mise en place. No dependencies. No side effects. If Treasury changes a parameter in its version tomorrow, Regulatory Reporting never notices. Because they don't share a bowl. They share a pantry.
That's what software engineering calls Open/Closed: open for extension (you can add new components), closed for modification (you don't touch the ones that already work). And in a kitchen it clicks instantly: you add bowls to the counter without touching the ones already there.
That's real decoupling. Not the kind written in an architecture document. The kind built into the structure of the data.
And this is where the pantry makes it possible
Remember how in Week 3 we insisted on separating identity from state? On every state snapshot of every entity being independent?
That separation wasn't just for traceability. It was for this.
When the pantry separates each entity's state snapshots, each prep in the mise en place can operate on a single entity without coupling to the others. The average balance operates on the account's state snapshots. Net premium operates on the contract's snapshots. The portfolio operates on the branch office's snapshots. Each one in its own bowl. Independent.
If tomorrow Risk adds a new indicator to its entity, Marketing's mise en place never finds out. If Finance changes the formula for technical provisions, Treasury's average balance keeps working. Because they don't share a chain. They share identities (the standard plugs from Week 1), but each one has its own mise en place.
In the domino-chain anti-pattern, a change in extraction breaks everything downstream. In a decoupled mise en place, a change to one entity only affects that entity's preps. The blast radius goes from “everything” to “only what it touches.”
Now think about what this means for the 15 teams from Week 1. Each team can create and evolve its preps without stepping on each other's toes, without coordinating, without waiting for someone else to finish. Exactly like 15 cooks prepping their bowls in the same kitchen without getting in each other's way.
What almost nobody does: a governed mise en place
In a restaurant, every prep in the mise en place has a spec sheet. It's not just a bowl of sauce. It's “Bolognese sauce — recipe #14 — Chef: Marco — Portions: 12 — Shelf life: 24h.” If Marco leaves and Ana has to make it, the sheet has everything.
In most data warehouses, transformations have no spec sheet. They're scripts in Git (if you're lucky) or in someone's personal folder. They have no explicit contract: what they expect to receive, what they guarantee to produce. No owner. No versioning.
A governed mise en place is something else entirely. Every prep has a declared contract:
- What it expects to receive: a time series with these columns, in this format.
- What it guarantees to produce: a series with no gaps, filled using this policy.
- Who defined it: the Treasury team, because they know how an average balance is calculated in their context.
- What parameters it accepts: granularity, fill policy, period.
- Which version it is: v1.0, approved on February 12.
That contract isn't a separate document. It's the component itself. Definition and execution are the same thing, just like the Recipe from Week 2. You can't run a prep without defining its contract. Governance comes built in.
In an amateur kitchen, the chef keeps the recipe in her head. In a professional kitchen, the recipe is on the spec sheet. And if the chef leaves, the sheet stays.
Back to our average balance
Remember the banking project from the beginning? 200,000 accounts, 15 indicators, 6 quarters. A 400-line script only one person understood.
With a decoupled, governed mise en place, the story would have been very different.
Gap filling would have been a standard component. Input: time series with gaps. Output: complete series. Policy: repeat last known value. That's it. Not 400 lines — a 10-parameter definition.
The average balance would have been another component. Input: complete time series. Output: average balance per period. Parameters: granularity, start date, end date. Another 10-parameter definition.
Two components. Composed. Each with its own contract. Each reusable across all 200,000 accounts and 15 indicators. Without rewriting a thing.
And when the regulator changed the definition from “business days” to “calendar days,” the change was one parameter. Not 200 lines of SQL. One parameter.
But the story has an epilogue that says more than the numbers. That senior engineer who wrote the 400 lines left the project three months later. It took the team two weeks to understand his code. Two weeks of a full-time engineer just deciphering what his predecessor had done. With a governed component, the contract would have explained everything: what it does, how it does it, what it expects, what it produces. Anyone could have made the switch from “business days” to “calendar days” in an afternoon.
Now imagine someone — or something — could take that parameter definition and automatically generate all the SQL needed for every warehouse, every source, every indicator. That what used to be a handcrafted 400-line script became a declarative definition executed by an engine.
As if someone had pre-cut every vegetable exactly the way you need it, before you even walked into the kitchen.
The numbers
For those who need to justify this to a steering committee:
Without a standardized mise en place (handcrafted, coupled scripts): - Time to implement a new average balance: 2-4 weeks (write the script, debug, validate with the business, document it if you're lucky). - Impact of a change at the source: unpredictable (it might break 0 or 15 transformations, and you won't know until Monday). - Investigation time when something fails: hours to days (where's the logic? who wrote it? which version is it?). - Reuse across sources: zero (every team reimplements the same calculation its own way).
With a decoupled, governed mise en place (components with contracts): - Implementation time: hours (you define the parameters, the component does the rest). - Impact of a change at the source: contained (it only affects that entity's preps; everything else never notices). - Investigation time: minutes (the contract tells you what it does, who defined it, and which version it is). - Reuse: total (the same average-balance component for 200,000 accounts; only the parameters change).
And there's one number that matters more than all the others: the number of different implementations of the same calculation. In a company with 15 domains and no standard components, it's common to find 8 or 10 versions of the “average balance,” each with subtle variations that nobody knows are intentional or bugs. With governed components, there's one version. Versioned. Auditable. And if someone needs a variation, they create a new component — they don't modify the existing one.
For the CDO or CTO reading this: think about what it means when the regulator asks “how do you calculate the average balance?” and the honest answer is “depends which team you ask.” That's not a technical problem. It's a regulatory risk. It's an audit finding. And with governed components, the answer is: “It's calculated by the SaldoMedio v2.1 component, with these parameters, defined by the Treasury team, approved on February 12.” One answer. Not ten.
Mise en place isn't a technical detail. It's the difference between a kitchen that scales and one that collapses every Monday.
What's next: plating
Today we prepped the ingredients. We saw that having organized raw data isn't enough: you need to turn it into business information. And we saw that how you do it matters as much as the result: independent functions, not coupled chains. With contracts, not loose scripts. Governed, not improvised.
But the prepped ingredients are still in the kitchen. The business doesn't walk into the kitchen to help itself. It needs a finished dish: a clean table, a dashboard, a regulatory report. Something it can consume without knowing how it was prepared.
Next week we move on to plating: where prepared information is presented for consumption. Where the complexity of the pantry and the mise en place is hidden behind a simple interface. And where we certify that whatever leaves the kitchen meets quality standards before it reaches the table.
Because a flawless mise en place is worthless if the dish goes out poorly plated. Or worse: if it goes out without anyone checking that it's fit to serve.
But before we move on to plating: