Your AI didn't find the cause. It found your definition.
A business metric I own looked like it fell 19.6% month over month. It had not moved.
I was at my desk trying to work out how to explain a rate that kept falling, so I asked our analysis agent what happened. A few minutes later it came back with a detailed attribution report, conclusion and supporting data included: 92.6% of the decline came from one user segment. Specific number, named segment, no hedging. I had the shape of the story and I was about twenty minutes from writing the doc that would have sent three people chasing a problem that did not exist.
What stopped me was not a process. It was that the number could not guide any decision. Our monthly business review looks only at creators who actually went live, while this rate was counted against everyone. The two did not line up, and the thing I had not checked was the denominator.
The model was not wrong
I want to get this part out of the way, because "the AI hallucinated" is the wrong lesson and it lets everybody off the hook.
The metric was a reach rate. The version in the dashboard divided by the whole population. But a large share of that population had not gone live that month, and the monthly review does not look at them at all. Counting them had nothing to do with whether the product was healthy.
That month, the mix shifted. More creators who had not gone live entered the population. The numerator held. The denominator grew. The rate fell.
So when I asked which segment drove the drop, the agent gave me the segment whose share had grown. That is the correct answer. The arithmetic was right, the attribution was right, and it answered exactly the question I asked on exactly the data I pointed it at.
The answer was still useless, because the question had a bad assumption inside it and neither of us checked.
Recompute against only the creators who actually went live and the same six weeks read 91 to 96%. Flat. A little boring. True.
Every rate has an eligibility rule, and almost nobody writes it down
Pick any rate metric you look at weekly. Conversion, activation, reply rate, resolution rate, retention. Each one is a fraction, and each denominator carries a claim about who was even in a position to produce the numerator.
That claim is usually not written anywhere. It lives in the head of whoever built the pipeline, and it was probably correct on the day they built it. Then the population changes underneath it. A new market opens. A gating rule ships. A segment grows. The definition does not update, because definitions do not have owners the way dashboards have owners.
You end up with a number that is computed correctly, monitored carefully, reviewed weekly, and quietly measuring something nobody meant to measure.
I have started calling this an ineligible denominator, mostly so I have something short to say in review meetings. The check that catches it is three words:
Who could have?
For whatever rate you are looking at: who was actually in a position to produce this outcome? Is that the group in the denominator? If the answer is "I think so," it is not.
It takes about a minute when the pipeline is documented and about an afternoon when it is not. The afternoon version is the one worth doing, because that is the case where nobody has looked in a while.
Why this gets worse with AI, and it is not the reason people usually give
The obvious worry is that models make things up. That is a real problem and it is not this one. This one is more annoying, because the failure has none of the warning signs we have been trained to look for.
Speed removes the pause. When a query takes an analyst two days, there is slack in the process. Someone reads it on the way. Someone asks a dumb question in a thread. When the answer lands in a few minutes, the pause is gone, and the pause was doing more work than anyone credited it for.
Confidence removes the tell. A junior analyst hands you an answer with a hedge attached, because they are unsure and it shows. The agent hands you 92.6% with no hedge, because it is not unsure. It has no way to be unsure about a definition it was never asked to evaluate.
And volume changes the shape of the mistake. One wrong analysis is an incident. A wrong definition wired into something that runs on a schedule is not an incident, it is a slow drift in what your whole team believes about the product, and there is no moment where an alarm goes off.
None of that is a model problem. The model did what it was asked. The gap is that no human had signed their name to the sentence "here is what right means for this number," and once nobody owns that, faster tools just get you to the wrong place sooner.
What this doesn't prove
Some things I want to be honest about, since it would be easy to over-read a single story.
It does not prove the 19.6% was meaningless. Against a different question, it is a real and interesting number. If you want to know how the overall population is shifting, that is exactly the metric to look at. It just was not the question anyone was asking that week.
It does not prove AI attribution is a bad idea. I still use it constantly and it is much better than what I did before. The agent found the segment faster than I would have, and once the denominator was fixed it re-ran in minutes. Speed is not the problem. Unowned definitions are the problem, and they were there long before any of this.
It does not prove I have solved anything. I caught this because I happened to be working out how to explain that number to the business, and I knew which group the review looks at, not because I had a check in place. I have one now: we rebuilt the metric around creators who went live, and I turned the "who is this counting?" check into a reusable instruction pack for AI and put it on our internal AI platform for operations, so anyone there can use it without knowing the trap exists. I do not know yet how many other numbers I look at every week have the same hole in them, and I am fairly sure the answer is not zero.
And one story is one story. I have not gone and measured how often this happens across a bunch of teams. If you have, I would like to read it.
The part that is actually good news
The fix here is not a framework and not a tool. Nobody has to build anything. Somebody has to answer one question in writing, once per metric, and put their name on it.
That is a strange thing to be optimistic about. But most of the AI failures I have run into at work have this shape. The model is fine. The pipeline is fine. There is a definition sitting in the middle of it that everyone assumed somebody else had checked, and the tooling got fast enough that the assumption now propagates before anyone notices.
So before you hand a metric to something that answers in a few minutes, ask who could have. It is not a big ask. It is the smallest version of owning the thing.
I work on evaluation for AI products that non-technical operators actually use. Next up: what happened when we A/B tested one of these agents against revenue, and the three ways I nearly got that number wrong too.