AI is not ready to answer questions about your data

Every other week, there’s a new story about AI solving a math problem that took mathematicians decades to touch. Most recently, it was a conjecture that had sat unsolved since 1939, cracked with an assist from an AI model. Meanwhile, the typical team still can’t ask ChatGPT what we sold last week without pulling three dashboards and arguing over whose number is right. That begs the question: why is AI intelligent enough to outsmart MIT professors, but it can’t outsmart the sales intern?

Here’s part of the answer. AI doesn’t get to guide a real business decision until it can answer with real accuracy, not 95%, not 99%, all the way. Getting there means clearing three hurdles: context, determinism and cost. Solve those three and you get accuracy. Right now, none of them are solved, which is why accuracy, not intelligence, is the actual blocker nobody wants to admit.

Grounding is the first wall

Ask an AI model a question about your company and it doesn’t actually know your company. It doesn’t know who owns which decision. It doesn’t know why your warehouse manager overrides the forecast every October, or which of two conflicting reports your team trusts. A new hire picks this up by working somewhere long enough. AI needs it handed over, deliberately, in layers.

The first layer is your teams, roles and workflows: who does what, and who is allowed to change it. The second is your industry and business context; the reason a service outage means something different for a bank than it does for a media company. The third is your data itself, the tables, definitions and history that make an answer true rather than plausible. All three have to be handed over deliberately. None of them show up on their own.

What I have observed is that companies skip straight to buying an agent and skip this grounding work entirely, which doesn’t end well. The agent still sounds confident. It’s also wrong in ways nobody catches until a decision has already been made on top of it.

Businesses do not want a coin flip

A client once told me something I have not stopped thinking about. We proposed solving their facility allocation problem with a classic optimization algorithm, deterministic and auditable, the kind where the same input always produces the same output. They were disappointed, stating, “I hired you because you are AI experts. Classic optimization gives us deterministic results. We want AI that gives us indeterminate results.”

I understood what they meant. They wanted something that felt like AI. But indeterminate is not a feature you want in a system telling you how much inventory to order. In the pursuit of using AI for the sake of using AI, it is easy to lose sight of the outcome the business actually needs.

That conversation still shapes how I scope every new engagement. When a client asks for AI without naming the decision it needs to support, I ask what happens the day it’s wrong. If the answer is a bad number in a board deck, we’re not talking about the same kind of AI they think they’re asking for. Sometimes that conversation ends the engagement before it starts. More often, it reshapes it into something smaller and more useful, an agent that handles the easy 80% of questions and flags the rest for a human, instead of one system trying to do everything at once.

Here is a number to think about. Anthropic recently published how it automated its own internal business analytics queries using Claude, and even with a purpose-built system, aggregate accuracy landed around 95%. That is one of the best AI labs in the world, building for its own internal use, still getting roughly one answer in twenty wrong. Tell any CFO that number and watch how fast they walk back to their deterministic BI dashboard.

I don’t think 95% is close enough. Not when the number ends up in a board deck. Not when a wrong answer becomes a real decision with real money behind it. In my experience, executives will forgive AI for being slow to learn. They will not forgive it for being confidently wrong. AI doesn’t earn a seat at the decision table by being right most of the time. It earns it by being right every time, the same way the deterministic system it’s replacing was. That’s not a popular thing to say in a market excited about what AI can do, but popularity was never the bar a business decision needed to clear.

Cost swings before the ROI math holds still

The variance in the cost of running an AI agent is enormous. I have watched the same agent perform the same task use up to 30 times as many tokens in one run as in another. Try defending a Well Architected Framework, or any ROI model, against a cost that swings that hard.

The trend lines pull in opposite directions at the same time. The price per token has generally been falling, which should make agents cheaper over time. At the same time, multi-agent architectures are burning through more tokens to do the same job. And an agent genuinely grounded in your teams, industry context and data — the grounding I described above — will use even more tokens than a shallow one. Full accuracy costs more, not less. The good news, if the trend holds, is that this improves rather than worsens over time. But “if the trend holds” is doing a lot of work in that sentence.

Not every question your business asks needs the same level of AI maturity. How many customers are in the system is easy, and I’ll trust an agent’s answer. How revenue has changed over time is medium. Forecasting future demand by product, week and store gets harder. Figuring out how to allocate demand across factories and manufacturing lines is super hard. Crafting a strategy that takes into consideration the three layers of context plus macroeconomic trends is, honestly, impossible for AI or a human to answer with certainty. The higher you climb that ladder, the more a wrong answer costs you, and the less I’m willing to accept anything short of fully right.

The mistake I see most often is a company jumping from easy straight to hard, expecting an agent that can answer “how many customers do we have” to also handle causal questions like “why did we stop selling a product.” Context requirements, accuracy requirements and token cost all climb together as you move up that ladder. That’s also the difference between handing an agent a task and handing it a role; the judgment a planner, buyer or analyst brings to a job every day sits at the top of the ladder, not the bottom.

The mathematician who cracked that 87-year-old conjecture and the CFO who wants last week’s sales are asking for different things. The mathematician wanted a collaborator: someone to try a thousand wrong paths and surface one interesting idea, where a 5% hit rate is a triumph. The CFO wants a number she can put in front of a board, where a 5% miss rate is a liability. AI has earned the first seat. It hasn’t earned the second.

It will, but not by getting smarter. It gets there when someone does the unglamorous work of grounding it in the company’s context, constraining it to be right the same way twice, and paying for that at a cost the ROI math can survive. Intelligence was never the blocker. Accuracy is. So, until an agent can clear all three hurdles, give it the easy 80% of the ladder, keep a human on the rest and don’t let the fact that it disproved a conjecture convince you it’s ready to order your inventory.