Last week we looked at the Recipe: who writes it, in what language, and why it changes the rules of the game. We ended by saying that a good recipe without ingredients is useless.
So let’s talk about the ingredients.
But first, let me tell you about something I lived through on a project that changed the way I think about how data is stored.
“Show me the original data from March 15”
We were on a project where the model had been in production for months. Everything worked. The reports reconciled. The users were happy. Until an audit review came along.
The auditor asked for something that seemed simple: “I want to see the original data, exactly as it came from the source system, on March 15. Untransformed. Uncleaned. The raw data.”
And the room went silent.
Because we didn’t have it.
The pipeline took the data from the source system, transformed it and loaded it into the model. The original data was overwritten with every load. What sat in the data warehouse was the result of the transformation, not the original ingredient.
“Okay, so can you ask the source system to resend the March 15 extract?”
More silence. The source system didn’t keep a history of extracts. What it had was the current state. The March 15 data, as it was on March 15, no longer existed anywhere. Not in the source, not in the warehouse. It was gone.
You know what we did? Rebuilt it. By hand. Cross-referencing load logs, partial backups of the source system, and the memory of a DBA who “seemed to recall” there had been an issue with the exchange rates that day. It took us three weeks. Three weeks to answer a question that, with the right architecture, takes three minutes.
And the worst part wasn’t the time. The worst part was the doubt. When we finally delivered the “reconstructed” data, nobody could guarantee 100% that it was exactly what had arrived that day. It was an approximation. A “we think it was this.” In front of an auditor, that doesn’t cut it.
But the story doesn’t end there. Two months later, the auditor asked for another data point from another date. And back to square one. Same team, same three weeks, same uncertainty. Because the problem wasn’t that specific March 15 data. The problem was that the architecture wasn’t designed to keep the original data. Every time someone asked “what was in here a month ago?”, the answer was “we don’t know for sure.”
And this isn’t an extreme case. Talk to any data team at a company of a certain size and ask them: “If tomorrow I asked you for the raw data that came in from the contracts system on the 15th of last month, would you have it?” Most will say no. Or they’ll say “it depends.” And that “it depends” is a regulatory risk many companies are carrying without even realizing it.
That’s when I learned something fundamental: the most important data in your data warehouse is the data you don’t touch.
The kitchen that throws away its packing slips
Back to our metaphor.
Imagine a restaurant receiving a delivery from its supplier: crates of vegetables, meat, fish. The cook opens the crates, grabs the ingredients and starts cooking right away. The crates, the packing slips, the labels with the packing date, the lot number, the supplier… all of it goes in the trash. “We’ve got the ingredients anyway, why keep the paperwork?”
Everything’s fine. The dishes go out. The customers are happy.
Until a customer gets food poisoning.
The health inspector shows up and asks: “Where did the chicken you used yesterday come from? Which farm? Which lot? When did it arrive? At what temperature was it stored?”
And the restaurant can’t answer. They threw away the packing slips. They didn’t write down the lot numbers. They don’t know which supplier yesterday’s chicken came from, because they buy from three different ones and labeled nothing.
This isn’t an exaggerated metaphor. In the European Union, food traceability is a legal obligation (Regulation (EC) No 178/2002). Every restaurant must be able to trace any ingredient back to its origin: supplier, lot, date of receipt. If you can’t, the inspector shuts you down.
Do you know what DORA, BCBS 239 and the EU AI Act require for data? Exactly the same thing.
Full traceability from the figure in the final report all the way back to the original data in the source system. Every transformation documented. Every origin identified. Every date recorded.
DORA, which came into force in January 2025, requires financial institutions to prove they can trace any data point in a regulatory report back to its source system. Not “we think it comes from here.” But “this data came from system X, on day Y, in batch Z, and this is the transformation applied to it.” If you can’t prove that, you have an audit finding. And audit findings in the financial sector have real consequences: remediation plans, regulatory follow-up and, in serious cases, penalties.
And it’s not just the financial sector. The EU AI Act is going to require exactly the same for any data that feeds an artificial intelligence model. If you can’t prove the provenance and quality of the data you train your model on, you can’t deploy that model in the EU.
And most “data-centric” companies are cooking straight off the delivery truck. No pantry. No packing slips. No labels.
Your data warehouse is the restaurant that throws away its packing slips. Everything’s fine until the health inspector calls.
What the Pantry is (and why it isn’t just a temporary holding area)
There’s a very common mistake when designing a data warehouse: creating a “staging” zone — a temporary area where data lands from the source system — and treating it as if it were the pantry.
It isn’t.
Staging is the loading dock. It’s where the supplier’s truck drops off the crates. You open them, inspect them and move them into the pantry. Afterwards, the dock gets cleared for the next delivery. If tomorrow you need something from yesterday’s delivery, it’s no longer on the dock. It’s gone.
The pantry is something else. It’s permanent. Every delivery is recorded as a batch with its date, its origin and its exact contents. What arrived, when, from where, and in what condition. If three months from now you need to know what came in on March 15, you open the pantry, look up the March 15 batch, and there’s the complete record.
And this is where the Shared Kernel from Week 1 really comes into its own. The business entities — Customer, Product, Broker, Branch Office — exist in the shared catalog independently of any source. They’re the company’s common language. The pantry is organized around those entities: each one has its own zone, its own shelf.
The Recipe from Week 2 is what connects each source to that common language. It says: “In this table, this field corresponds to the catalog’s Customer entity, and this other one to Product.” It’s the mapping between the source system’s language and the company’s shared language. Thanks to that mapping, the pantry knows where to put everything that comes in.
But — and this is what separates a good pantry from a junk drawer — the ingredients don’t just get piled in a heap. They get organized.
Think of a real professional pantry. The vegetables aren’t mixed in with the meat. The dairy isn’t next to the fish. Each type of ingredient has its zone. And every ingredient has its label.
In a data warehouse, this organization follows a fundamental logic: separating what a piece of data IS from what a piece of data HAS at a given moment.
What a piece of data IS is its identity.
A customer is always the same customer, no matter when you look or which system they come from. A product is always the same product. A broker is always the same broker. These identities are stable. They’re registered once in the pantry. If the same customer comes in from another system under another field name, the pantry recognizes them as the same person thanks to the Recipe’s mapping. It doesn’t duplicate them. It integrates them.
What a piece of data HAS is its state at a specific moment.
The same customer may live in one city in January and another in June. Their risk profile may change from one quarter to the next. Each change is stored as a dated snapshot. Earlier snapshots don’t disappear: they coexist. You can see what that customer looked like at any point in time.
This seems like a minor detail, but it has enormous consequences. Remember the audit from the beginning? “Show me the data from March 15.” If every state snapshot is dated and coexists with the earlier ones, the answer is immediate: you open the March 15 snapshot and there it is.
And then there are the relationships: what connects to what.
This customer bought this product through this broker at this branch office. If tomorrow they buy another product, that’s another relationship. The first one is still there. Relationships are the glue between identities.
In the source system, all of this arrives mixed into a single row of a flat table. Identities, states and relationships scrambled together across 200 fields. But in the pantry, it gets separated. Everything in its place. And that separation is what changes everything.
Why does this separation matter so much?
Because it solves three problems in one go:
First, traceability.
Want to know what address the customer had on March 15? You don’t need to reconstruct anything. You look up the March 15 state snapshot and there it is. No archaeology. No “we think it was this.”
Second, surgical reprocessing.
You discover a bug in a transformation that’s been in production for three weeks. You need to regenerate the data from March 1. Without a pantry, this is a nightmare: you have to ask the source system to resend the data. If it even can. If the data from three weeks ago hasn’t already been purged. With the pantry, it’s all there. You fix the transformation, rerun it on the pantry data, and within hours you have corrected reports.
You don’t throw out the whole fridge: you just pull Tuesday’s tray.
This seems like a minor nuance, but in practice it’s huge. Without a pantry, a typical reprocess means: calling the systems team, requesting a re-extract, waiting for them to schedule it, checking that the data is the same as what originally arrived (which you can’t guarantee), reloading, retransforming and praying there are no side effects. I’ve seen reprocesses take two weeks just to get the source system to send the data again. With the pantry, you open the batch, rerun, and you’re done. It’s the difference between putting out a fire with a garden hose and putting it out with a fire extinguisher.
Third, parallelism. Remember the 15 teams from Week 1 that needed to work without stepping on each other’s toes? This separation is what makes it possible. If identity is separate from state, Risk and Marketing load their attributes for the same customer in parallel. Each one adds its own state snapshots to the same identity. No conflicts. No waiting. And the regulatory project from Week 2, where IT played detective hunting for “Product” in every system? With the pantry, Product is already sitting on its identity shelf, with all the mappings declared in the recipes. Nothing to guess.
“Who changed the format of this column without telling anyone?”
There’s a problem every data engineer has lived through at least once.
One Monday morning, the pipeline that has been running without a hitch for six months fails. The error: a field that used to be numeric now arrives as text. Or a column that used to be 8 characters wide is now 10. Or a field that always had a value now comes in empty for 40% of the records.
What happened?
Someone in the source system made a change. A migration, an upgrade, a “minor tweak” nobody communicated to the data team. Because in many companies, the teams running the source systems don’t even know there’s a data warehouse downstream that depends on their data.
And now the detective work begins. Who made the change? When? Was it intentional or a bug? Does it affect all records or only new ones? Since what date has the data been wrong?
If your pantry is a temporary loading dock, these questions are an ordeal. You can’t compare “how it used to arrive” with “how it arrives now” because you didn’t keep the “before.”
But if you have a permanent pantry, the answer is immediate:
- When did it change? You open the pantry and compare yesterday’s batch with today’s. The field changed format between Friday’s load and Monday’s.
- Does it affect earlier data? No. Earlier batches are intact in the pantry. You can verify that last week’s data is still correct.
- Who do I ask? And this is the piece almost nobody has.
The supplier’s label: who owns each ingredient
In a professional restaurant, every crate that arrives at the pantry carries the supplier’s label. If the chicken smells off, you don’t call the chef. You call the supplier. Directly. No middlemen.
In a traditional data warehouse, that label doesn’t exist. Data reaches the central team and nobody knows who’s responsible for what. What does this field mean? Why did the format change? Is this null correct? Every question goes to the central team, which doesn’t know the business meaning of each piece of data. And so begins the chain of emails, meetings and weeks of investigation.
But if you’ve been following the logic of the previous weeks, the answer is already there.
In Week 2 we said we empower the business user. We give them tools. We let them be the one who defines the recipe for their data: which entities it contains, how they’re mapped, which rules apply. We give them the power to build their own data products.
But as a certain superhero’s uncle once said: with great power comes great responsibility.
If you define the recipe for your data, you own that data. Not because someone assigned you an “owner” field on a governance form. But because you defined it, you mapped it, you loaded it. It’s yours. When something breaks, the question comes to you. Because you’re the one who knows.
Data ownership isn’t an administrative act. It’s the natural consequence of empowering the person who knows the data.
And in practice, this changes everything:
- A field changed format → the pantry knows that source was loaded by the Product team (because they defined the recipe) → the question goes straight to them.
- A field has 40% unexpected nulls → the pantry knows that source belongs to the Customers team → the question goes to them.
- A table stopped arriving → the pantry knows Operations is responsible → you call them.
Without going through IT. Without a mass email to 15 people. Without weeks of investigation.
When the pipeline doesn’t fail… but the data does
Let me tell you what happened one real Monday on one of my projects. The nightly pipeline loaded the weekend’s data without errors. No process failed. No alert went off. All green.
But on Tuesday, a user opened their report and the numbers didn’t add up. A key metric had dropped 35% overnight. Impossible.
We investigated. The problem: the source system had changed the format of a numeric field on Friday afternoon. A field that used to come in as “1234.56” now arrived as “1.234,56” — the thousands and decimal separators had been swapped. The data loaded, got transformed, all without a single technical error. The pipeline did exactly what it was told. But the result was garbage, because the ingredient had gone bad and nobody inspected it on arrival.
Do you know how long the investigation took? Two days. Two days to discover that the problem wasn’t in the transformation, or the business logic, or the model. It was in the ingredient.
The pipeline didn’t fail. It did exactly what it was told: cook with a spoiled ingredient.
If someone had compared Friday’s delivery with Thursday’s and said “heads up, this field has a different pattern,” we would have caught it before loading. In minutes, not days.
That’s exactly what a permanent pantry makes possible: because you have all the earlier batches, you can compare. No history, no baseline. No baseline, no inspection possible. The pantry doesn’t just store — it’s the foundation quality is built on.
How that inspector works, exactly what it detects and how it certifies the data before it reaches the plate… we’ll get to that later, when we talk about plating.
The numbers
For those who need to justify it to a committee:
Without a pantry (temporary staging, overwritten data): - Average time to answer an audit on historical data: 2–4 weeks (manual reconstruction, partial backups, “we think it was this”). - Cost of a reprocess: days to weeks (you depend on the source system to resend the data, if it even can, if anyone even knows how to run a historical extract). - Incidents from uncommunicated changes at the source: discovered in production (when the report is already wrong), slow resolution (who made the change? when? who do I ask?). - Ability to work in parallel: limited (everyone shares the same tables, one team’s changes can step on another’s).
With a pantry (permanent, organized, with owners): - Average time to answer an audit: minutes (the original data is right there, immutable, with a date and a batch). - Cost of a reprocess: hours (you rerun the transformation on the pantry data, without depending on anyone). - Incidents from changes at the source: detected at load time (before they affect the model), direct resolution (the owner is declared, the alert goes to the person who can fix it). - Ability to work in parallel: full (identities are shared, each team adds its own attributes without conflicts).
Think about what this means for a company with 200 source tables and a regulator that asks for more data, faster, with more traceability, every year. Without a pantry, every audit is a crisis. Every reprocess is a negotiation with the systems team. Every change at the source is a ticking time bomb that can go off in production weeks later.
With a pantry, the audit is routine. Reprocessing is a button. The change at the source is caught before it does any damage. And the 15 domains work in parallel without stepping on each other’s toes.
And if you run a data team and you’re thinking “okay, but how much does it cost to maintain a permanent pantry?”, the answer is counterintuitive: the pantry is cheaper than not having one. A terabyte in the cloud costs between €20 and €25 a month. A single badly answered audit — three weeks of a 4-person team reconstructing data — costs between €15,000 and €25,000 in working hours. Add the reprocesses, the incidents caught too late, and the loss of the business’s trust every time the data team says “we don’t know for sure.”
The pantry isn’t an expense. It’s insurance. And like any good insurance, you hope you never need it. But the day you do, it’s worth a hundred times what it cost.
What’s more, the pantry enables something that’s impossible without it: real automation. If the raw data is organized with separate identities and declared owners, 80% of the work your data engineers do by hand today (investigating incidents, rebuilding history, hunting for who to ask) disappears. Your team stops putting out fires and starts building value.
The pantry isn’t a technical luxury. It’s the foundation everything else is built on. Without it, you’re cooking straight off the delivery truck and throwing away the packing slips. It works right up until it doesn’t.
What’s next: the mise en place
Today we opened the pantry. We saw why storing data permanently and in an organized way is the most important decision you can make in a data architecture. And we saw that organizing the pantry means separating what a piece of data IS from what it HAS, so that every team can work in parallel without stepping on each other’s toes.
But the ingredients in the pantry, organized as they are, are still raw. Unprocessed. No business rules applied. The tomato is on the vegetable shelf, labeled and dated. But nobody has washed, peeled, chopped or seasoned it.
Next week we step into the mise en place: where raw ingredients are turned into business information. Where the calculation rules (“Net Premium = Gross Premium − Ceded Reinsurance”), the derivations and the aggregations are applied. Where raw data becomes something the business can use to make decisions.
And we’ll see why, if the pantry is well organized, the mise en place is surprisingly simple. And if it isn’t… well, you know how that goes.
But before we get into the mise en place: