VitrX Limited is delighted to announce the acquisition of Wentworth.

This brings together two highly established and respected IT businesses with a combined experience of nearly 50 years,

Illustration of a wall-mounted meter with a red gauge dial spilling out an endless printed cost ledger onto the floor, with an AI robot working at a desk in the background, representing the rising cost of running AI agents at scale.

The most expensive words in AI: ‘Let’s just put an agent on it’

AI agents are becoming easier to access, while the cost of running them at scale can climb quickly. A stronger measure of value starts with the completed business outcome, including the full cost of producing it.

“Let’s just put an AI agent on it.”

It is easy to see why that idea is gaining ground. AI agents can do far more than respond to a prompt. Give them a goal and they can plan a sequence of steps, use tools, consult data, check their work and carry a process through to completion with limited human input.

That capability creates real opportunities for productivity. It also creates a new cost problem, because every extra action can consume more context, compute, integrations and paid services.

A July 2026 paper from QuantumBlack, AI by McKinsey, examines this through the lens of “agentic economics”: the real cost and value of work performed by AI agents. For business and IT leaders, the useful question is whether the reliable outcome produced by an agent is worth everything required to deliver it.

Cheaper intelligence can still produce a bigger bill

The headline economics of AI look remarkable. The Stanford HAI AI Index cited by McKinsey reported that the inference cost of GPT-3.5-level capability fell from $20 per million tokens to $0.07 through 2024.

Enterprise spending moved sharply in the other direction. Menlo Ventures research cited in the paper found that enterprise large language model spending tripled over a 12-month period by the end of 2025. A separate May 2026 Enterprise AI FinOps Survey cited by McKinsey found that 93% of 75 qualified respondents had exceeded their AI budgets. McKinsey’s forthcoming 2026 State of AI survey also found that one in five respondents had constrained AI use because of AI-related operating costs.

Those studies use different samples and methodologies, so they should be read individually. Together they point to a clear operational challenge: lower unit costs can encourage far greater consumption.

The same pattern has appeared elsewhere in technology. Storage became cheaper and businesses stored more. Cloud made compute easier to access and organisations learned quickly that convenient capacity still needed financial controls. AI is developing along a similar path, with usage expanding at machine speed.

The McKinsey paper captures the idea in one useful line: “Tokens are not value; tokens are the bill.” The value sits in the completed outcome, such as a resolved support case, a qualified opportunity, an accurate analysis or a correctly processed claim.

Why agents can consume so much more than chatbots

A chatbot usually responds to a request. An agent works through a goal, and that autonomy creates activity. It may retrieve records, call applications, compare answers, repeat failed steps, ask another model to review its work and carry a growing amount of context from one stage to the next.

McKinsey highlights six drivers that can overwhelm the benefit of falling model prices. Three stand out immediately.

First, context accumulates. Large language models are stateless, so instructions, history and documents often have to be supplied again as a workflow progresses. Research cited in the paper found that some agentic coding tasks can consume roughly 1,000 times more tokens than simpler chat or single-turn reasoning tasks.

Second, refinement can become the expensive part of the process. Checking, repairing and reverifying an answer consumes additional resources. One study cited by McKinsey estimates that around 60% of the cost of an agentic software-engineering task can sit in the refinement loop.

Third, autonomy introduces variability. Two runs of the same task may take different routes, call different tools or make different numbers of attempts. Research cited in the paper found a factor-of-30 variation in completion cost for the same programming task. Average-cost assumptions become much less useful when the expensive tail is significant.

The other drivers are equally practical. Organisations may use their most capable and expensive models for simple work. Orchestration choices can add cost without improving the final result. Information structure also matters, because prompt design, context length, formatting and data structure influence token consumption.

There are costs outside the model too. Security gateways, monitoring, data retrieval, orchestration platforms and software integrations can all add to the bill while handling the same workflow. Agentic cost belongs to the whole system.

Measure the outcome, then measure the cost

Tokens, prompts and model calls are easy to count because the technology exposes them. Business value needs a more useful unit of measurement.

For customer service, that could be the cost per issue resolved correctly without reopening. For sales, it might be the cost per genuinely qualified opportunity. For finance, it could be the cost per reconciled account within an agreed error threshold.

This approach changes model selection too. A higher-cost model can be the economical option when it reaches the right answer quickly and reduces human rework. A cheaper model can become costly when its outputs require repeated checking or create downstream errors.

Human review, security controls, integrations and rework all belong in the calculation. The same outcome should be used when comparing machine work, human work, traditional automation and hybrid approaches.

Allocate intelligence according to the workload

McKinsey argues that businesses should allocate intelligence with the same discipline used for capital. Different workloads need different levels of capability.

Low-volume expert reasoning may justify a flagship model because accuracy and capability carry more weight than unit cost. High-volume, predictable classification may suit managed open or private models where scale economics matter. Regulated workloads may require private cloud or sovereign environments for residency and control. Very high-volume internal tasks may become candidates for private or on-premises deployment.

Most enterprises are likely to operate across several placement models at the same time. The valuable capability is the routing and governance layer that sends each workload to the right environment according to volume, predictability, sensitivity and quality requirements.

This is also where IT architecture and commercial management come together. Model choice, context management, caching, evaluation, security and workload placement all affect the economics of the final outcome.

Your business context may be the real differentiator

As foundation models become more widely available, access to the model itself becomes less distinctive. The quality of the context around it grows in importance.

Customer history, operational data, decision precedent, product knowledge, process telemetry and exception handling can all change the quality of an agent’s output. Two organisations can use the same commercial agent and achieve very different results because the information available to that agent is structured, governed and connected differently.

Data readiness therefore belongs inside the AI strategy. It affects performance, security, sourcing and the ability to change platforms later.

It also changes the build, buy and partner conversation. Work that creates or depends on valuable proprietary context may need to stay close to the organisation. More generic activity can be placed elsewhere when the business case supports it. The right answer will vary workload by workload.

Give every material agent a mandate, a budget and a stopping rule

Agent governance is an operating-model decision. Every material agent should have a defined purpose, an agreed outcome, appropriate access to systems and data, escalation rules and a clear performance measure.

It should also have a budget and a stopping rule. An autonomous system can retry, research and call additional tools repeatedly. That persistence needs a ceiling in any consumption-priced environment.

Ownership should be explicit as well. The CIO may own the information layer. The CTO may oversee architecture, routing and technical efficiency. The CFO may measure return and forecast spend. Operations leaders understand the process, exceptions and quality thresholds. Clear accountability helps prevent duplicate activity and unmanaged cost.

Questions worth answering before an agent scales

Before moving an agent from an impressive demonstration into everyday operations, leadership teams should be able to answer a short set of questions:

  • Which completed business outcome are we buying?
  • What is the full cost per successful outcome, including human review, integrations, security, monitoring and rework?
  • How much does cost vary between runs, and what happens at the expensive tail?
  • What level of accuracy, latency and autonomy does the process genuinely require?
  • Which tasks need frontier intelligence, and which can use smaller models or conventional automation?
  • What proprietary context improves the result, and how is that data governed?
  • Where should the workload run given its volume, predictability, sensitivity and quality threshold?
  • Who owns performance, budget, exceptions and the stopping rule?
  • What evidence would cause us to expand, redesign or stop the deployment?

These questions create the conditions for sustained adoption. They make cost, value and accountability visible before usage grows faster than the operating model around it.

From AI experimentation to machine-work management

The first phase of enterprise AI encouraged access and experimentation. Teams needed room to explore the technology and find useful applications. The next phase requires a management discipline around machine work.

That means making the economics visible, matching workloads to the right technology, treating context as a strategic asset and setting clear guardrails around autonomy and spend. Hybrid environments are likely to play a major role, spanning cloud services, private infrastructure, on-premises systems and AI-capable devices at the edge.

For VitrX, this is the conversation that matters when organisations move from AI trials into operational use. Start with the business outcome, understand the context that makes it valuable, decide where the workload should run, define the acceptable cost and risk, and keep enough flexibility to adapt as models, pricing and regulation change.

VitrX can support organisations in turning those decisions into a practical technology roadmap, helping to assess where workloads should run and plan the cloud, data centre, networking, security and infrastructure needed around them. The aim is to create an environment that can adapt as AI use grows while keeping performance, resilience, control and cost visible.

An AI agent may be capable of doing the job. Long-term value depends on whether the operating model around it allows that job to be delivered reliably, securely and economically at scale.

“Let’s just put an agent on it” can be a useful starting point. The next questions determine whether it becomes a sound investment.


References

QuantumBlack, AI by McKinsey.
“Is that AI agent worth it? Agentic economics and the modern operating model.” July 2026.

Stanford Institute for Human-Centered AI.
“AI Index 2025 Annual Report.”

Menlo Ventures.
“2025: The State of Generative AI in the Enterprise.” December 2025.

Bai et al.
“How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks.” April 2026.

Salim et al.
“Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering.” January 2026.

The May 2026 Enterprise AI FinOps Survey and McKinsey’s forthcoming 2026 State of AI survey are referenced within the McKinsey article above.

Share With:

Recent News

Charity is for life, not just at Christmas

Dr. Seuss once wrote: “What if Christmas, he thought, doesn’t come from a store? What if Christmas, perhaps, means a […]

Disruptive or conformity – which is your path for the future?

Which describes your current server and SAN setup? Are you just expanding your existing estate? Have you considered a new […]

New Aged Celebration or Legacy Panic?

With the newly launched Intel Kablake processors comes a bitter pill. Faster more efficient systems balanced with the inability to […]