- Blog
- AI Agent Costs: Why You Pay for the Attempt and Not the Result
AI Agent Costs: Why You Pay for the Attempt and Not the Result
There is a line item on your AI invoice that nobody prints. Call it the Retry Tax. It is the money you spend on attempts that failed. It is also the hours you sat waiting while they failed. Almost every AI agent pricing model in the market today charges you for both. You pay for the tokens the agent burned. The tokens make the answer worse. You pay again for the retry. Nobody writes that number down. So nobody manages it. And 73% of enterprises told the FinOps Foundation their AI costs blew past budget in 2026.
Think about how strange this is. A plumber who cracks your sink does not invoice you for the second visit. A lawyer who files the wrong motion does not bill you for the refiling. Software is the only trade that convinced its customers to pay for the swing rather than the hit. We agreed to it without arguing. For thirty years, software did the same thing every time you ran it. A spreadsheet does not have a success rate.
An agent does.
AI agent pricing bills the attempt, not the result
Look at what is actually published. Intercom’s Fin charges $0.99 per resolved outcome. If it does not resolve the issue, you are not billed. Zendesk prices in the same neighborhood, roughly a dollar to two dollars per automated resolution. Salesforce’s Agentforce charges $2.00 per conversation, and that charge applies whether or not the issue gets resolved.
Same category. Same buyer. Opposite arrangement.
Run the numbers on ten thousand conversations a month. Fin reports an average resolution rate around 76%. At $0.99 per outcome, you pay for the 7,600 that landed, which comes to $7,524. At $2.00 per conversation, you pay for all ten thousand, which comes to $20,000.
That is 2.66 times the cost for the same volume of work.
Sit with the second half. Of the ten thousand conversations, 2,400 went unresolved. Under per-conversation pricing you paid $2.00 for each of them anyway. That is $4,800 a month, roughly $57,600 a year, for work the machine did not do. You bought an attempt. The attempt failed. The invoice did not notice.
None of this makes Salesforce a villain. Agentforce is a serious product and the company has been visibly wrestling with how to price it. It launched per conversation, added a credit model at roughly ten cents per action, then added per-user licensing at $125 a seat. Three models running at once, changed twice inside eight months. In June it agreed to buy Fin for about $3.6 billion, which is one way of saying the outcome-priced model was worth billions to own.
That is not a company failing. That is a company discovering that the meter is hard to defend.
The tokens make it worse, which is the part that stings
Here is where the Retry Tax stops being an annoyance and starts being a design problem.
Anthropic published a detailed writeup of its multi-agent research system. Buried in it is the most useful number in the industry. Its agents used roughly four times the tokens of a chat. Its multi-agent setup used about fifteen times. And token usage alone explained 80% of the variance in performance on one of its evaluations.
Read that plainly. These systems work in large part because they spend. In the same post, Anthropic said something else. Work that needs shared context, or heavy coordination between agents, is a poor fit today. They disclosed the cost, they disclosed the benefit, and they named the boundary. That is more honesty than most vendors offer, and it deserves saying out loud.
But if spending is the lever, then the buyer is holding the wrong end of it.
Now add the finding almost nobody has absorbed. The research team at Chroma tested eighteen frontier models. That included the ones selling million-token context windows. They measured how accuracy holds up as input grows. All eighteen degraded. Not some. Every one. Effective context landed somewhere around half to two thirds of what was printed on the box. Strangest of all, the models did better on shuffled text than on well-organized documents.
Put the two findings side by side and the sequence appears.
You pay per token. More tokens raise the odds of a good answer, up to a point. Past that point, more tokens make the answer worse. The agent produces something wrong. You retry. The retry costs tokens too. The failed attempt now sits in the context. So the second run starts from a worse place than the first.
You paid three times. Once for the failure. Once for the degradation. Once for the fix.
The bill nobody forecast
The most quoted defense of all this is that inference is getting cheaper, so the problem solves itself. Prices did fall. One analysis of enterprise API traffic put the blended cost per million tokens down about 67% year over year. And yet 73% of enterprises still exceeded their AI budgets.
Total spend is price times volume. Prices fell. Volume did not. Agentic architectures multiply volume on purpose, because volume is the thing that makes them work.
Uber burned its entire 2026 generative AI allocation in four months. Five thousand engineers, adoption climbing from 32% to 84% in a single month, power users clearing two thousand dollars each. Then a $1,500 cap. Tesla, which once ran leaderboards ranking engineers by token consumption, now caps spend at $200 a week.
These are not careless companies. They learned that agentic spend cannot be forecast, because the thing that drives it is failure, and nobody forecasts failure.
The villain is the arrangement, not the vendor
It would be lazy to say the model providers designed this to enrich themselves. There is no evidence for it, and the accusation is beneath the argument. Anthropic published the token multiplier itself. Nobody hid anything.
The problem is structural. When a vendor is paid for consumption, a wrong answer costs the vendor nothing. When a vendor is paid for outcomes, a wrong answer costs the vendor everything.
Those two structures produce different engineering. Not different morals. Under the meter, the Retry Tax is your problem. Under outcome pricing, it becomes the vendor’s problem. Vendors solve their own problems fast.
The answer is to price the work, not the attempt
Nobody hires a contractor who refuses to quote the job. No ceiling, no refund for the cabinet installed backwards, billed by the swing of the hammer. Yet that is the standard arrangement with agentic AI, and we signed it without reading it.
There is a better one. It is what we believe the industry is moving toward, and it is the premise we are building Paciva on.
Quote the job. Bill the result. Tell the buyer the price before the work starts.
Under that arrangement the tokens become the vendor’s cost of goods, not a line on your invoice. Failed attempts come out of the vendor’s margin. Context rot becomes an engineering problem for the people who can actually fix it. The buyer stops financing the machine’s mistakes and starts buying finished work.
To be clear about where we stand: this is our thesis about where digital labor is going, not a product announcement. We are not publishing a price today.
The market is already voting. Seat-based pricing fell from 21% to 15% of companies in twelve months. Gartner expects at least 40% of enterprise SaaS spend to move toward usage, agent, or outcome-based pricing by 2030. Fin, which does not bill you when it fails, sold for billions.
Why this matters more to enterprise than to anyone else
A founder with a corporate card can absorb a surprise. A procurement organization cannot.
Look at what actually kills enterprise AI projects. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, and the causes it names are escalating costs, unclear business value, and inadequate risk controls. Model quality does not appear on that list. Forrester found 56% of organizations see no measurable financial benefit from their AI spend.
Now put outcome pricing against those three causes, one at a time.
Escalating cost stops escalating, because the price is agreed before the work begins. Unclear business value clarifies itself, because the invoice is denominated in outcomes rather than in tokens, and an outcome is a thing a CFO can recognize. Risk controls become tractable, because a system that must produce a result in order to get paid is a system with a reason to stop and ask before it does something irreversible.
That last point is the one people miss. Governed autonomy and outcome pricing are the same idea wearing different clothes. If you are paid per token, an agent that plows ahead and gets it wrong is a revenue event. If you are paid per outcome, that agent is a loss. Only 6% of organizations say they trust agents to run a core process end to end. Under the meter, nobody has an incentive to fix that number.
There is a procurement argument here too, and it is unglamorous and decisive. The metered token bill is one of nine cost buckets, and eight of them are unmetered. Retrieval, storage, orchestration, the engineer who babysits the run, the analyst who checks the output. A finance team can forecast a quoted job. It cannot forecast a meter attached to a system whose consumption rises when it fails.
That is why the share of FinOps practitioners managing AI spend went from 31% to 98% in one year. The enterprise did not get more curious about AI. It got scared of the invoice.
