The unit is doing a lot of work
Salesforce now has a billing unit for agent work, and whoever signs your renewal should read the definition twice. diginomica reports that Salesforce defines an Agentic Work Unit as "one discrete task accomplished by an AI agent" where "raw intelligence is converted into real work". An AWU covers a processed prompt, a completed reasoning chain, or a tool invocation.
Those are not the same size. A processed prompt is one round trip. A completed reasoning chain can fan out into a long sequence of steps. A tool invocation fires every time the agent reaches into a system of record. One business outcome, resolving a case or drafting a quote, might land as a handful of units or several dozen depending on the route the agent picked. The customer asked for the outcome. The invoice counts the route.
The scale figures in the same report explain the timing. diginomica cites Agentforce annual recurring revenue above $1.5 billion in Q2, up 240% year over year, with 7 billion AWUs delivered to date and 3.2 billion of those in Q2 alone. Salesforce has consumed over 19 trillion tokens. Close to half the lifetime unit count arrived in a single quarter, which tells you what these curves do once agents get past pilot.
The sentence a budget owner should read twice
Patrick Stokes, President of Applications and Marketing at Salesforce, described the consumption risk to diginomica: "One engineer can generate like $100,000 a month bill without much problem." The vendor said that about its own metering model. We'd take the sentence straight into the next budget review.
Marc Benioff, CEO of Salesforce, frames the value around "what our platform does with it" rather than the token itself. That's a reasonable position for a company selling outcomes. It also means the number of billable units sitting between an employee's request and a finished piece of work is a product decision Salesforce makes and your finance team inherits. The source doesn't give a unit price or a bundled allotment, and that's the first gap we'd want closed in writing.
What to ask about consumption forecasting
Go into the renewal conversation with scenarios rather than seat counts. Pick the five agent use cases you expect to run at the highest volume and ask the account team what one run of each consumes in AWUs, as a measured number rather than a modeled one. Then ask what the count does when a reasoning chain fails halfway, and when a tool call times out and the agent retries down another path. The source says nothing about failed or abandoned work, and that silence is where surprise invoices live.
Ask about variance as well. An average is no use if the tail is fat. If five percent of runs consume ten times the median, the forecast has to be built from the tail. Ask too whether consumption reporting arrives in near real time or only after the period closes, because a monthly report tells you what you already spent. Anyone who has worked through a licensing change checklist knows reporting granularity is most of the negotiation.
Keep the commercial terms in the same conversation as the technical ones. The usual discipline for a multi-year renewal applies harder with a consumption metric, because neither side has a historical baseline and the vendor owns the model.
Chargeback has to exist before agents go wide
If AWU spend lands in one central IT line, every team has an incentive to build agents and none to build efficient ones. The chargeback model needs to be running before your second or third agent ships. Retrofitting attribution after a quarter that blew through budget is a far worse project than setting it up now.
At minimum, every agent carries an owning cost center from the day it is created, and every unit of consumption can be traced to the agent that spent it and the team that asked for the work. That's a metadata problem more than a finance problem, and it sits alongside the questions raised by governance across Agentforce editions. If you can't say who owns an agent, you can't say who pays for it.
The same dynamic shows up in the total cost of a Salesforce edition. The license line gets negotiated hard once a year. The consumption line grows quietly in between.
What a spend cap actually attaches to
A cap on the org total does almost nothing useful. It fires after the money is gone and it punishes whoever happened to run last. Thresholds belong further down, on an individual agent and on the team that owns it, with a separate ceiling on what any single run is allowed to consume.
The per run ceiling is the one people skip. An agent that loops or works its way through a long chain of tool calls can spend heavily with nobody watching, and long-running agent work is where that shows up first. A hard stop on units per run converts a runaway into an incident ticket instead of an invoice line.
Where the cap gets enforced matters as much as the number on it. If your agents reach models and tools through a shared control point, policy gets set once, which is the argument behind runtime policies at an AI gateway. If every team wires its own path, you're left with the vendor console and whatever alerting it offers, which the source doesn't describe.
The question to bring to your next meeting
Salesforce has handed the market a unit and a set of very large numbers attached to it. What nobody has yet, customers included, is a defensible way to forecast next year's consumption from this year's pilot.
So put this on the steering agenda. If the agents already running in production tripled their volume tomorrow, could anyone in the room say within twenty percent what that costs? If answering means calling the account team, the chargeback model and the per run cap both need to land before the next agent goes live.



