A deflected case is three different outcomes
Every service automation business case rests on a deflection number, and that number is usually measuring something much looser than the people quoting it believe. Deflection gets recorded whenever a self-service session ends without a case being created, which quietly bundles three outcomes that have nothing in common.
The first is a customer who asked a question, got a correct answer and went away satisfied. The second is a customer who gave up, closed the tab and decided the problem was not worth the effort. The third is a customer who found the answer and then raised a case anyway about something unrelated. Only the first one saved you money.
The second outcome is a cost you have already incurred and have not yet seen. It surfaces as churn at renewal, as a complaint escalated through an account manager, or as a second contact on a channel nobody joins up with the first. Counting it as a saving means the programme books a benefit for producing a loss.
The cheapest work is the first to go
The unit that matters is fully loaded cost per contact rather than an agent salary. Fully loaded means pay plus recruitment, training, quality assurance, supervision, licences, telephony, workspace and the senior time spent on cases that bounce back. Business cases built on salary alone understate the real cost per contact and then overstate what removing a contact returns.
Here is an illustrative example, with round numbers picked to show the shape of the problem rather than to describe any real service desk. Say password resets and order status checks are half the contacts and a fifth of the handling cost, because they finish in two minutes while a billing dispute takes forty. Automate every one of them perfectly and the queue halves while spend falls by about a fifth.
What remains is the expensive residue. The cases left behind need more context, more judgement and more time from someone senior enough to make a decision without asking. The work that survives automation costs more per case than the old average did, and it needs people you cannot hire in a month.
Average handle time rises when the programme works
Follow that arithmetic into the quarterly operations review and it turns uncomfortable. Average handle time across the remaining queue goes up, because the short contacts left and the long ones stayed. The team performed exactly as the design intended and the headline metric moved the wrong way.
A service organisation measured on average handle time will therefore look worse after a successful deflection programme. Managers who understand the mechanism start protecting the metric instead, which tends to mean discouraging the automation in small ways, or reclassifying work so that short contacts reappear in the denominator.
Name the trap before the programme starts. Agree with finance and with service leadership that handle time will rise, agree roughly how far, and agree what you will watch in its place. Arguments about a metric nobody can reproduce twice burn more goodwill than a genuinely bad quarter ever does.
The saving moves to a line that scales with traffic
Consumption pricing changes the shape of the cost rather than removing it. Headcount is a fixed cost you can plan a year ahead, while agent actions charged per interaction or per resolution form a variable line that rises with traffic, including the traffic that fails. A badly grounded agent spends money producing a wrong answer, and then the case arrives anyway and you pay for the same contact a second time. Anyone sizing this should start from what Agentforce actually is at the metering level.
Model that failure mode before you sign anything. Work out what the programme costs during a volume spike, during a bad release, and in the fortnight after a knowledge change sends containment down. The pricing question behind agentic work units is the same question in different vocabulary, and the answer belongs in the business case rather than in a renewal conversation eighteen months later.
Content maintenance is the other standing cost. Deflection quality is knowledge quality, and knowledge decays every time a price changes, a policy moves or a release renames a screen. Somebody has to own the articles, retire the stale ones and rewrite the ones that generate a follow-up contact every time they are read. Budget that role or watch containment drift down quarter by quarter.
Escalation design decides what a failure costs
Answer quality gets the attention in the demo, and escalation design determines what happens on the day the answer is wrong. Every automated conversation fails some of the time. The recoverable failure hands the customer to a person with the transcript, the account, the entitlement and the thing they were trying to do already attached.
The unrecoverable failure drops the customer into a queue where they explain everything again to somebody who cannot see what the agent already told them. That contact now costs more than it would have if you had never automated it, and the customer remembers the repetition rather than the resolution.
Look closely at how the handover writes into the case record. If the agent conversation lives in a separate object that nobody joins to the case, the escalation path loses its context by design. Getting the case data model right matters more here than another round of prompt tuning, and routing configuration decides whether the handover lands with someone who can finish the work.
Measure resolution rather than session endings
Measure resolution instead of absence. Ask whether the problem was solved, using an explicit signal such as a confirmation step at the end of the conversation, and record silence as unknown. An unknown can be investigated. A false success walks straight into the benefits column and stays there.
Follow the person rather than the session. Take a window of seven or fourteen days, take every customer who touched self-service, and check whether they came back through any channel including phone, email, the partner portal and the account manager inbox. Sessions are easy to count and they hide the behaviour you most need to see.
Hold a control group if your volumes allow it. Route a small random share of eligible traffic straight to a person, keep it there for a full quarter, and compare resolution, repeat contact and satisfaction across both populations. Nothing else separates the effect of the automation from the effect of a quiet month, a product fix or a seasonal dip. If you cannot run one, say so in the business case rather than presenting a before and after as proof.
So when somebody puts a deflection number in front of you, ask one thing before anything else. Of the sessions counted as deflected, how many can you show contacted us again within fourteen days through any channel at all. If the reporting cannot join channels, that number describes how often people stopped trying, and the saving you were promised is still sitting somewhere in next year's contact volume.



