The morning the chats stopped arriving

A support team I worked with spent a Tuesday morning watching the chat backlog grow while six agents sat in Omni-Channel with nothing to do. Nobody had touched a routing setting that week. The agents were online and the work would not move.

The cause was a case routing configuration charging one unit per case against a presence configuration granting five units. Overnight the case queue had pushed five cases to each person. Everyone was full on paper, so chat, which shared the same pool, had nowhere to go.

Routing failures usually read that way. The configuration does what it was told, and what it was told stopped matching how the team works. Aim for a setup whose behaviour at the busiest hour you can predict.

Channels, configurations and queues each do one job

Four records carry routing in Service Cloud. A service channel connects one Salesforce object to routing, so cases move through one channel and messaging sessions through another. A routing configuration holds the routing model, the priority and what a work item costs in capacity. A queue holds the work and names who is eligible.

The join between them is the part people miss. Each queue points at one routing configuration, so the rules travel with the queue rather than the channel. Two queues on the same case channel can route in completely different ways, which confuses everyone when nobody wrote down which queue uses which.

The fourth record is presence. A presence configuration sets a person's capacity, whether work is accepted automatically, and whether they may decline. Presence statuses are separate records linking a person to the channels they can receive work from. Draw all four before you open setup, because most later arguments are about that drawing.

Capacity decides how much work a person can hold

Capacity governs everything you will observe. The routing model picks who gets the next item, and capacity decides whether anyone is eligible at all. A short queue hides that. A long queue turns capacity into the behaviour your service manager complains about.

The default model counts units. A presence configuration grants a person a number of units, and each routing configuration says what an item of that channel costs, as a count or a percentage of their total. Weighting is the point, so a live conversation should cost more than a case someone picks up after lunch.

Status-based capacity moves the limit onto the presence status instead. Someone on a chat status can hold three conversations, and on a case status fifteen cases. That model tells the truth when people switch between kinds of work across a shift, and the unit model tells the truth when they hold several kinds at once.

Ask the team how many live conversations a good agent runs before quality drops, and how many open cases they carry alongside. Those two numbers are your configuration. If nobody can answer, start low and raise it after two weeks of watching.

When two routing configurations disagree

Priority sits on the routing configuration and a lower number wins. When a person has spare capacity and several queues hold waiting work, Omni-Channel pushes the highest priority item first, and equal priorities fall back to whichever item waited longest.

What gets reported as broken routing is almost always this. Chat sits at priority one, cases sit at priority five, and the chat queue never empties in business hours. The case queue grows all day and nothing is technically wrong, because everyone is busy with higher priority work. It is the drift that makes lead routing rules stop making sense two years in.

The fix is a decision rather than a setting. Either some people are reserved for case work through their own queue and presence status, or case priority rises with age through an escalation, or the case service level was never real. Pick one with the service manager, and write down which queue starves on a bad day.

Online and available are different states

An agent who is online in Omni-Channel may still be unavailable for your channel. Presence statuses carry the channels they cover, so a person on a messaging status never receives a case however much capacity they have. Plenty of missing-work tickets end right there.

Work already pushed to a person stays with that person. Going offline does not hand cases back to the queue; they keep that agent's name until someone closes, transfers or reassigns them, the practical side of a queue not being an owner. Live conversations go worse, because a customer waits while the session goes unanswered. Closing the browser is no shift handover.

Push time-out on the routing configuration looks minor in the form and saves you here. With auto-accept switched off, an item sits on an agent's screen until they accept it, and with no time-out it sits there all shift. Set the time-out in seconds, decide whether an expiry moves that agent to a status that stops new work, and confirm the item is offered to somebody else.

Skills routing fails closed

Skills-based routing is the most capable part of Omni-Channel and the easiest to break quietly. Work reaches only people holding the skills an item requires, which gives real control over language and tier. It fails closed, so an item needing a skill nobody holds any more routes nowhere.

Nothing flags that item as failed. It waits in a queue with no eligible agent, and waiting work looks calm on a dashboard that counts errors. Build skills from a list someone maintains, confirm two active users hold every skill you route on, and set a skills time-out so requirements relax after a few minutes.

Overflow is the other half of the same problem. A route work action in an Omni-Channel flow takes a fallback queue, and that queue catches what the rules could not place. A fallback queue with no members does nothing, so give it real people, a status covering the channel, and someone who checks it each morning. Deflection economics push the easy contacts away, so what reaches skills routing is harder.

Before routing goes live, walk this list with the service manager and whoever owns it after you move on. Each line is something you can check inside an hour.

  • Every queue names the routing configuration it uses.
  • Capacity per person comes from a number the team agreed.
  • Each channel's capacity cost matches the attention that work takes.
  • Priorities are deliberate and someone named the queue that starves.
  • Push time-out is set and expired work goes to another agent.
  • Two active users hold every skill you route on.
  • The fallback queue has real members who come online.
  • A supervisor can see work still waiting and who is full.

Test with two users and one saturated agent

Routing cannot be verified by one admin clicking around. With a single user and a single item everything routes, because the interesting behaviour appears only when one person is full and another is not. Two logged-in users is the minimum, plus somebody in Omni Supervisor.

Run the saturated case on purpose. Fill one agent to capacity, leave the second on a status that does not cover the channel, and push work in. The correct outcome is work that waits rather than work landing on someone who cannot take it. Then free one unit and watch the item move.

Break things next. Have an agent let a push expire, and confirm the item reaches somebody else and the status changed as designed. Have an agent go offline holding two cases, then see where those cases are ten minutes later. Route an item requiring a skill nobody holds, and time how long before anyone notices. How you model the case record shapes what the supervisor sees.

Do the skills test before go-live rather than after, and keep the supervisor view on a screen someone watches for the first week. When somebody says routing is broken, ask who was at capacity that minute, then which routing configuration that queue was using.