Seat counts were knowable and meters are not

Consumption-priced AI capability produces a renewal surprise for a structural reason, and that reason has very little to do with how any particular vendor behaves. Seat licensing has one genuinely useful property. Cost is knowable in advance because headcount is knowable in advance. A platform owner who can count the people needing access can predict next year's bill within a rounding error, without reading anything technical and without asking anyone's permission.

Consumption pricing breaks that link on purpose. Spend is now set by how much the system gets used, and usage is set by thousands of small decisions made by people who have never seen a rate card. A developer adds a grounding step to improve an answer. An admin widens a trigger so a flow catches more records. A business owner extends an agent to a second region because the first one went well. Every one of those decisions is defensible on its own terms, and not one of them passed through procurement.

So the organisation has moved the commercial decision out of procurement and into engineering and operations without telling either. Procurement still owns the contract and still believes it owns the spend. Engineering owns every choice that actually sets the number and mostly does not know that it does. People who follow licensing changes closely tend to spot the gap first, because they are the only ones reading both the agreement and the release notes.

The four shapes a consumption surprise takes

Pilots account for most of the surprises. A capability arrives free or heavily discounted while the vendor wants adoption, a team builds something useful on top of it, and the discount ends on a date nobody wrote into the architecture decision. The technical work was sound. The commercial assumption underneath it expired quietly, which is the same pattern behind cost control problems after a pilot on any platform with a free tier.

Bundling produces a larger version of the same problem. Where a capability is sold inside a higher tier rather than separately, the question stops being whether to buy a feature for one team and becomes which tier the whole estate sits on. That is a much bigger commitment, and it usually surfaces late, because the feature was what everyone discussed and the tier was what everyone assumed.

The third shape catches architects specifically. Usage types get named before they get priced. A vendor documents that some action counts toward a billable unit, teams start designing against that meter, and the rate lands later. An organisation can spend a year building a dependency on a unit whose price does not yet exist. What is fixed and what the vendor can set later varies by agreement, so read yours rather than trusting what a colleague remembers from a different contract.

The fourth shape is growth that does not track headcount. An automated process runs as often as something triggers it, and a rollout that goes well increases triggering. Hire nobody and the bill still climbs, because the automation you shipped last quarter now runs against three more record types in two more regions. Automation owners who lose the thread usually lose it at exactly this point.

Nobody owns the number in the middle

Every one of those shapes gets worse when nobody owns the number. The platform team sees technical telemetry, calls and runs and durations, and reads it for health rather than for cost. Procurement sees an invoice with a handful of lines on it. Neither sees consumption broken down by the team that caused it, and that breakdown is the only view letting anyone act while the total can still be changed.

Non-production environments consume quietly. Test runs, regression suites, a developer iterating on a prompt all afternoon, a load test somebody scheduled nightly and forgot about. Whether non-production usage bills at the full rate, a reduced rate, or nothing at all depends entirely on your agreement, so check your own terms before you assume the sandbox is free.

Retries and failures usually cost the same as successes. An agent that is poorly grounded, gets an answer wrong, gets corrected and tries again has billed twice for one unit of value. A flow that hits a malformed record and retries five times has billed five times for nothing. The engineering fix and the cost fix are the same fix here, which is one of the few conveniences in this whole area.

Attribution is the governance work that pays

You cannot manage what you cannot attribute, so mapping consumption back to owning teams is the highest-value governance work available. It is also unglamorous. It means tagging environments, naming service accounts after the team owning them rather than the project that created them, and keeping a register of which agent, flow or integration belongs to which cost centre.

Most platforms give you enough to do this if you decide to. Environment-level separation is the crude version and it works. A deliberate environment strategy that puts each business unit's automation in its own environment turns one untraceable total into several attributable ones, and that is most of the work. Where finer telemetry exists, use it, and do not wait for perfect attribution before you start with coarse attribution.

Capture a baseline while you still can. Usage stops being freely observable the moment it starts being charged, because the first thing that happens when a bill appears is that somebody throttles or blocks whatever is generating it. Ninety days of ordinary usage recorded before a pricing change is worth more in a renewal conversation than any amount of reconstruction afterwards.

Showback changes what teams build

A team that sees its own number behaves differently. Showback does not have to mean chargeback, and in most organisations it should not, at least not in the first year. Sending each team a monthly figure for its own consumption, with last month's figure beside it, changes design choices without anyone writing a policy. Engineers who know a retry costs money write better error handling.

Name one owner for the total. That person has to be technical enough to read consumption data and work out what produced it, and senior enough to stop something growing faster than anyone planned for. Those two requirements often point at different people, which is why the role goes unfilled so often. Give it to the platform owner and make sure they talk directly to whoever signs the renewal.

The owner's job between renewals is boring and specific. Watch the trend by team, ask about anything that doubles, and keep a written record of what changed and why it changed. That record is what turns the renewal conversation into a negotiation rather than an explanation.

Ask about the rate before you depend on it

Consumption pricing is not a trick. It genuinely lets a team start small without a procurement cycle, which used to be the single biggest source of delay in getting anything useful built. Somebody can test an idea in an afternoon on a capability that would once have needed a business case, a budget line and a three-month wait. The same flexibility making adoption easy makes drift easy, and no vendor invented that trade-off to catch anyone out.

The negotiating point follows straight from the mechanism. The time to ask about a rate is before you have built a dependency on it. A question about future pricing asked during a pilot, when you can still pick a different design, gets a different answer from the same question asked at renewal, when the vendor can see the integration running in your production tenant. Negotiating a multi-year renewal goes better when you brought the usage data rather than received it.

Pick the three highest-consuming workloads you can identify this week, find out which team owns each one, and send those teams their own ninety-day numbers with nothing else attached. Do it before the next pricing change, while the numbers are still easy to pull and nobody has a reason to make them smaller.