Connect with us

NEWS

Cheaper Tokens Made AI Deployments Cost Enterprises More

Gartner’s inference paradox shows agent workflows costing over five times a chatbot even as tokens cheapen, while most firms still cannot see the spend.

Published

on

Gartner said on August 17, 2026 that AI inference costs per agentic workflow will rise more than fivefold through 2028, even as token prices keep falling. That split is now the working math behind expensive AI deployments, from software invoices to megawatts.

Enterprises spent the first half of 2026 pushing models to do more work per prompt, and finance teams met the bill after the usage had already landed.

Falling Token Prices Raised Total AI Costs

Will Sommer, a senior director analyst at Gartner, calls the split the inference paradox: better unit economics that still raise the total cost of AI, without a clean path to matching value. Routing a task to an agentic reasoning model, Gartner found, raises provider inference costs by at least five times versus a basic chatbot, and often more as the job gets harder.

The same research group has said agentic systems use 5 to 30 times more tokens per task than a standard chatbot. Gartner still forecasts that inference on a one-trillion-parameter model will cost providers over 90% less than in 2025 by 2030. The cheaper unit is real. The workflow is hungrier.

THE THREE FORCES GARTNER FLAGS

  • Cheaper foundations: Base model cost economics are improving quickly, which is the slide finance teams keep quoting.
  • Richer models: Those savings get spent on more powerful, more expensive models as soon as the cheaper ones make them affordable.
  • Longer jobs: Agents reason, retry, and talk to other agents, so they burn far more tokens than a single chatbot reply.

Gartner now puts the agent-workflow increase at more than fivefold through 2028. Product teams that treat falling token prices as a green light will scale the expensive path by default.

Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens. There is no reliable, economical one-size-fits-all model on the horizon. Producing competitive AI products will require developing and maintaining complex multimodel ecosystems.

Will Sommer, senior director analyst, Gartner, August 17, 2026

Sommer’s warning is practical. Defaulting to generic autonomous intelligence, he said, produces unbounded costs orders of magnitude above a product that routes easy work to cheap models and holds the dear reasoning for the hard steps.

Tokenmaxxing Reached Finance After the Damage

Through early 2026 a lot of teams treated token volume as a proxy for progress. The habit even got a name, tokenmaxxing: stuffing context, looping agents, and letting copilots run because the unit price looked small on a dashboard.

Nicholas Merizzi, a principal at Deloitte Consulting, said the timing is what hurt. “Before CFOs could even get it on their radar, the bills, the damage had been done.” Sommer put the model-side cause in one line: “The models produced by leading AI labs are getting more token-hungry faster than they are getting cheaper.”

Chasing raw token volume is a bad score for the work. It rewards disposable scripts and a feeling of motion, and it is a score token vendors have little reason to replace with a price tied to business value. Several large employers already had to stop treating token counts as a performance metric after staff learned to climb the leaderboard.

Palantir, which sells deployment muscle rather than metered tokens, made that argument in public on July 1, 2026. The companies that still sell by the token have the opposite incentive, and finance only sees it when the invoice arrives.

Only 31% Can See AI Software Spend

The Flexera 2026 State of ITAM Report, released June 24, 2026 and based on 512 technology professionals, is the cleanest picture of that lag. AI now sits across cloud, SaaS, data, and devices, so it does not land in one ledger line.

THE FLEXERA 2026 VISIBILITY GAP

Measure Share of respondents
Accurate visibility into AI software 31%
Wasted AI spend up year over year 59%
Complete visibility across IT assets 36%
Audited in the last year 48%

Flexera also found that only 31% have accurate visibility even among shops already trying to track AI inside software spend. Microsoft still led audit activity, cited by 64% of respondents over three years.

Becky Trevino, Flexera’s chief product officer, said AI is changing the economics of IT faster than most organizations can adapt, a familiar pattern of rapid adoption and then a scramble for control as spend surges. A separate Harness report, circulating in July, found that 1 in 4 dollars spent on AI goes to waste and that more than half of businesses still lack a dedicated owner for those costs.

You cannot govern a bill you cannot see. That is why FinOps teams, not model labs, are being asked to invent a unit price after the fact.

The Linux Foundation Wants a Token Price List

The FinOps Foundation, a Linux Foundation program, published its sixth annual State of FinOps survey on February 19, 2026. Among 1,192 respondents, representing companies with more than $83 billion in annual cloud spend, 98% of FinOps teams manage AI spend, up from 31% two years earlier. Nine in 10 of those practitioners now manage SaaS as well.

J.R. Storment, executive director of the FinOps Foundation, said FinOps practices will be critical to C-level decisions about multi-year technology bets as AI costs rise. Among large spenders of $100 million or more, about 68% are already using or testing FOCUS-formatted billing data.

FOCUS is the FinOps Open Cost and Usage Specification, built so AWS, Microsoft, Google Cloud, Oracle, and others emit comparable cost files. The next fight is to make that file speak tokens, agents, unused seats, and GPU hours in the same breath as a virtual machine.

THE COST LINES STILL HARD TO MATCH

  • Tokens and credits: ChatGPT, Claude, Gemini, and Copilot bills that spike with retries and long context, not with headcount.
  • Seats that sit idle: Enterprise copilots bought in bulk while teams expense other models on personal cards.
  • GPUs and agents: Clusters, containers, and multi-step agents that do not map to a user license.

Trevino pointed at the FinOps and Tokenomics Foundations as the place that work is moving. Channel partners already talk about a services line in answering a question most CIOs still cannot: what does one AI feature cost to run.

CISOs Fear Agents They Cannot Inventory

Cost is only the invoice. The other unpaid line is permission. Okta’s Global CISO Insights 2026 survey of 306 security chiefs found 81% worry their AI systems are not properly governed. Only 47% said they can identify every agent in the environment, and 46% said they can control what those agents can reach.

Fewer than half of those CISOs believe their boards treat AI security as a business enabler. Security gets framed as a brake on a productivity program the CEO is already selling, so the controls arrive after the agents are in production.

Gartner’s Shiva Varma has argued against one-size-fits-all trust for agentic systems. Each agent and each process needs its own permissions, because a summarizer and a purchasing agent are not the same risk. That design work costs engineering time, which is another reason the bill does not fall when the token price does.

Shadow AI makes the inventory worse. If a business unit can paste customer data into a consumer chatbot, the official Copilot contract is not the spend that matters, and it is not the breach path either.

How Much Power Do Towns Have?

Local governments now price AI with permits, water, and substations. Data Center Watch counted at least 75 U.S. projects worth $130 billion disrupted by opposition in the first quarter of 2026, and a June Ipsos poll found only 14% of Americans would back a data center in their community. PJM can cut data-center load before blackouts under a May 19, 2026 Energy Department order.

THE 2026 GRID AND SITING FIGHT

  1. End of 2025: Trackers count 396 active local opposition groups around U.S. data center sites.
  2. First quarter of 2026: Those groups more than double to 833 across 49 states, and at least 75 projects worth $130 billion are blocked or delayed.
  3. May 19, 2026: The Energy Department lets PJM Interconnection curtail data centers that have backup generation as a last step before rolling blackouts.

The political math is ugly for builders. A town of a few hundred people can stall a gigawatt interconnect because the scarce input is deliverable power, not a welcome banner. Cedar Hill, Tennessee, population 304, put a moratorium on data centers on June 1, 2026, one of several small jurisdictions that decided the promised jobs did not cover the water and the rate shock.

Hyperscalers are already shopping for on-site generation so they are not waiting in the utility queue. That spend never shows up on a token invoice, but it is part of the same 2026 receipt: the model got cheaper per word, and the building that serves the model got harder to plug in.

Vendors Now Sell Engineers With the Model

If the software does not configure itself, someone has to sit in the customer’s office and wire it to payroll, inventory, and policy. Palantir named that person a forward-deployed engineer more than a decade ago. In 2026 the model vendors copied the job title, because the bottleneck moved from the model to the last mile.

LinkedIn said job listings for forward-deployed engineers rose twelvefold over the past year. Lightcast put median advertised pay at more than $188,000, against about $145,000 for ordinary software-engineer ads.

THE 2026 VENDOR DEPLOYMENT BETS

Vendor When What they stood up
OpenAI May 11, 2026 $4 billion Deployment Company, including Tomoro
AWS June 2026 $1 billion forward-deployed engineering organization
Microsoft July 2026 AI-deployment unit staffed with thousands of experts

Anthropic and Google Cloud made similar moves the same season. Palantir cofounder and CEO Alex Karp told investors the entire business nearly doubled in 12 months, and EVP and CTO Shyam Sankar used the earnings call to draw a line. “Remember, only Palantir has FDEs,” he said. “Everyone else has sparkling sales engineers.”

Channel firms that planned to train up on the next model are already late. Enterprises will not pause a rollout for a partner academy when the vendor will embed its own people this quarter.

With the pace of change as rapid as it is, FDEs are going to be critical in converting opportunities to revenue. Enterprises don’t really have the patience at the moment to wait for the partners to get trained up, because by the time they do, the newest model comes out.

Peter Bryant, GSI practice lead, Omdia

The 2026 receipt is not a single line item. It is a token invoice that arrived before the CFO had a forecast, a security program that cannot name its own agents, a substation fight in a town that never asked for a cluster, and a new class of engineers whose job is to make the model do something a seat license never did. Slowing down does not appear to be on offer. The people who can meter the work, and the people who can sit inside it, are the ones collecting.

Harry is the editor and lead writer of WEBWIZARD 360, which he owns and runs independently for readers around the world. Ten years in journalism, the early ones reporting and the later ones editing, shaped a simple rule about technology coverage: a vendor's claim stays a claim until it has been tested or documented. Benchmarks are run on the device itself, changelogs and filings are read in full, and a launch announcement is checked against what actually ships. He carries the same caution into the other nine sections, so business stories start with the accounts, science stories with the paper, and sports, entertainment, lifestyle, travel, auto, gaming and general news with whatever official record exists. Numbers are verified before publication, without exception. If an article turns out to be wrong, it is corrected on the page with a note that says what changed, in line with the corrections policy the site publishes. Reader mail reaches him at support@webwizard360.com, and he replies to it himself.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending