Skip to content
Article

AI spend is only as predictable as your controls

4 Aug 2026

We made the case in Beyond Per-Seat that consumption pricing in the age of AI moves variability from the vendor’s side of the table to yours. That variability requires governance; otherwise, leadership may be in for some rude surprises. 

Bain’s 2026 survey of 951 companies found that close to 40% of those measuring their AI savings came in under 10%, well short of the 11 to 20% they had targeted. In addition, 90% are raising their AI budgets again anyway. The reason for the shortfalls? Bain ties the gap to how AI is deployed and governed, not to the technology. The teams with predictable AI spend are the ones governing it at the action level, while it’s happening.  

The pricing model meters exactly what it’s built to meter. Governing what the AI does in production, and how much it consumes, falls to the organization.

Cheaper tokens won’t necessarily lower spend

It’s easy to think that, because model costs are dropping, spend will sort itself out. Fortune’s reporting complicates that notion. Even though inference on a large model is expected to get dramatically cheaper per token by the end of the decade, enterprise costs are still projected to rise. That’s because agentic workflows call models far more often per task than a typical chat assistant. In the agentic age, consumption grows faster than price falls. 

For a contact center running AI agents at production volume, waiting for prices to fall isn’t a strategy. A resolution that runs through several reasoning steps and a couple of tool calls already costs more than its so-called list price. And the gap widens as you automate more of the queue.  

Instead, what moves the numbers is how much work each interaction is allowed to generate. Those are guidelines that each individual organization can and should set themselves. 

Surprise invoices trace back to missing controls

Overage(s) tend to get filed under pricing, when the cause usually lives a layer down, in governance. It’s where the financial results of unsupervised AI decisions surface. A few of those decisions drive most of the variance, and the contact center has rarely had to own any of them: 

  • Scope: which interactions an agent takes on, and where it has to hand off.
  • Model selection: whether a routine balance inquiry quietly runs on a premium reasoning model that costs many times what the task needs.
  • Consumption per resolution: the model and tool calls an agent accumulates before it closes a case. 

Picture an AI agent that re-queries the model each time a tool returns, then hands to a second non-human agent that repeats the retrieval from scratch. The case resolves, the customer is satisfied, and that one interaction costs several times what a dashboard may imply. None of it shows up as a line item on the invoice, and all of it is yours to govern.

Cost predictability is a post-launch discipline

Other parts of the enterprise already manage consumption spend continuously rather than approving it once. The agentic model raises the stakes: the Bain report finds only about 7% of companies run fully autonomous agents, with the common setup still routing decisions to a person for approval. A business case built on full-automation economics rarely matches a system running with a human in the loop, and the spend follows what’s running rather than the projection. 

Cisco frames the same shift on the operations side as a question of trust rather than raw speed. Autonomy without guardrails scales a wrong decision as fast as a right one, so the answer is to put review and approval in front of an AI agent’s actions, and an audit-and-learn step behind them. What Cisco calls AgenticOps keeps AI detecting, recommending, and acting while governance and human oversight are embedded in the workflows like built-in checkpoints. 

It’s the same idea with regard to spend: before launch, you set what an agent can do and what it must escalate, you watch what it actually does, and then you revise on a cadence instead of at renewal. Cloud and engineering teams already use this rhythm to keep consumption from drifting.

Observable spend is governable spend

If you can’t see an agent’s cost behavior while it happens, the invoice is the first place that spend takes visible shape. By then, the decisions behind it are weeks old. 

This is where Cisco’s investment in observability reaches the cost conversation, and where security and visibility turn out to be one posture seen from two sides. The Webex roadmap now folds monitoring, governance, and security for AI agents into a single management layer. It follows an agent across its lifecycle, and pairs with quality evaluation that scores every interaction rather than a sample. 

The logic is plain enough. You can’t secure or cost-control behavior you can’t see, especially while it’s underway. That visibility has to span the systems an agent reaches into, not just the metered ones.

How Bucher + Suter builds the governance in

This is the work we do with CX and IT leaders before architecture and contracts lock the choices in. We define where agentic AI should and shouldn’t operate in your environment, and we set the rules for model selection and escalation that follow. Then we instrument the deployment so cost is observable while it’s being incurred. Finally, we tie that spend to the outcomes you’re already accountable for. 

Here’s what we’ve found: capping and fixing tends to suppress the value of the deployment, or push the cost somewhere less visible. Pricing purely on outcomes swings too far the other way, handing the system a reason to chase the metric ahead of the intent behind it.  

As Justin Greis noted in Computerworld, an AI rewarded for cutting service costs can learn to deflect legitimate contacts to hit the number. A more durable approach is to make spend observable and govern it deliberately. 

If you’re standing up agentic AI, or already watching the spend climb, let’s map where it’s heading before the next renewal. before the next renewal.