I'm not going to ask you to picture a scene. I'm going to tell you about one I lived through this very week.
I had a documentation system I'd spent months building: rules, processes, criteria, everything an agent needed to look up to know how to behave.
One day I had to make a big change. Something structural, not a tweak.
I went to make it, and found I couldn't.
That change touched so many places (a bit here, a nuance there, a cross-reference somewhere else) that isolating it was nearly impossible. Every document had too many concepts tangled up inside it, and pulling on one thread dragged all the others along with it.
I ran the numbers, and they didn't add up. It was cheaper to throw the whole thing out and start from scratch than to try fixing it piece by piece.
So that's what I did.
I rebuilt the entire system in a couple of hours, far less than it had taken the first time, because by then I knew which questions to ask.
But the fact that I had to scrap it all was still a sign that something in the original design was wrong.
That "something" is exactly what's sinking most generative AI projects in companies today. And almost nobody is paying attention to it.
The symptom everyone misdiagnoses
When an AI assistant fails in production, the easy diagnosis is always the same:
"The model is hallucinating."
And most of the time, that's not true. What's failing isn't the model. It's what you've been feeding it. And it fails because of two very different laws that almost never get told apart.
We shovel in gigabytes of uncurated PDFs and expect the model to work magic with them. And the result is the same as always: the model answers, with all the confidence in the world, based on something nobody has reviewed and nobody could change without breaking half a dozen other things.
With all the technical horsepower we have today, why do we keep tripping over the same rock?
The problem isn't the technology. The problem is that almost nobody is governing what gets fed to the model. We call it AI Governance, and today, in most companies, it doesn't exist.
Law 1: Diminishing returns
Economics has a principle that's more than a century old: the first worker you add to a factory increases output. So does the tenth.
But there comes a point where every new worker starts getting in the way, because the fixed resource (the floor space, the machines) doesn't grow with them. From then on, each additional worker doesn't add. They subtract.
Language models work exactly the same way. We throw everything in "just in case we need it someday." Volume grows without judgment because more information sounds like safety, never like risk. And the bigger it gets, the harder it is to govern.
We've been here before. In the Big Data days we called it the data swamp: a repository that starts out promising and ends up as a place where nobody can find anything. The name has changed (today it's a RAG, not a Data Lake) but the mechanism is identical.
Think of a kitchen pantry. At first it's tidy: spices on one shelf, cereals on another. The cook walks in, grabs what they need, and cooks.
Now imagine that for three years nobody has thrown anything away. Expired jars at the back. Three different brands of the same thing, all open, all half used. Two versions of the same product, an old one and a new one, and nobody knows which is the right one.
The cook walks in to make dinner. And even if they're the best cook in the world, it'll take them three times as long, and they might mix up the salt and the sugar because nothing is labeled.
You're not giving your AI a pantry. You're giving it a junk closet.
Law 2: Concept fragmentation
This one is subtler: it's not about how much volume there is, but about how it's spread around.
The problem isn't that documents are fragmented. The problem is that the same business concept shows up repeated and scattered across several places at once, and nobody has decided where its single home is (its Single Source of Truth).
Here's a real example, the kind you'll find in any midsize company. The word "Customer" is defined three different ways in three departments:
- Marketing defines it as anyone who has interacted with a campaign, whether they bought or not.
- Risk defines it as the holder of an active contract.
- Sales defines it as someone with recent activity.
All three definitions are correct in their own context. None of them is wrong. But if you ask the AI "how many customers do we have?", the model has three truths in front of it and no way to decide which one is right.
The model doesn't fail because it's dumb. It fails because you've put three contradictory truths in front of it and asked it to pick one. When a concept lives in three places at once, neither the model nor the most disciplined person on earth can tell which one is in charge.
The bill both of them run up
Here's the effect almost nobody sees coming: even as models get more powerful, the cost doesn't go down. It goes up.
Volume without judgment drives up cost directly: every extra token gets billed, no exceptions.
And concept fragmentation drives it up indirectly, which is even sneakier: the model ends up reading garbage that shouldn't even be within its reach (retired versions, duplicate definitions, out-of-scope content) and pays to process it just as it would for what actually matters.
But the financial cost isn't the worst part. The worst part is the cost in trust.
When someone asks the AI something that should have one clear answer, and the AI comes back with a number that doesn't match anyone else's, trust breaks. And once it's broken, it's very hard to rebuild.
Why we shouldn't call this "hallucination"
"Hallucination" is a convenient word because it suggests the model made something up out of thin air, as if it were running a fever. But that's not what's happening.
If you feed a model inconsistencies, ambiguities and dead versions, it isn't hallucinating on a whim: it's trying to reconcile several lies that you put in front of it.
"AI doesn't invent the chaos. It inherits it, and serves it back with the same confidence it would use to tell the truth."
Calling it "hallucination" is convenient. It lets us blame the model and keep stuffing garbage into the pantry. But it isn't honest, and above all, it isn't useful: you can't fix a problem you've misdiagnosed.
The trap: "let's just dump it all into a RAG and see what happens"
At this point, someone always says the magic words: "let's dump everything into a RAG (a search system for AI) and let the model figure it out."
It sounds great. It solves the problem in one stroke: you don't curate anything, you don't organize anything, you don't decide anything. You throw it all in and the model does the magic.
But when you put it into practice, an enemy worse than disorganization shows up: the false sense that the problem is solved. The system answers. It generates text. It seems to know. For the first few weeks, everyone is thrilled.
Until someone asks something that matters. Until the number the AI gives doesn't match the one from the official system. And then comes the awkward question: "where did it get that?"
And nobody knows.
What's next
These two laws explain why the problem exists. They don't yet explain how to solve it.
The good news is that we don't need to invent anything new. The data world has been wrestling with exactly this problem for thirty years (volume without judgment, information without governance) and it already has a mature answer. We just need to apply it to what we feed an LLM today.
That's what comes next.
If you look at the sources and documents in your own RAG today, can you find what you're looking for on the first try? What makes you think a model will do any better than you?