The pilot worked, and that proves less than you think
The request usually arrives as a forwarded email. A team built an agent in Copilot Studio that answers policy questions from a SharePoint library, or an Agentforce agent that drafts replies for one support queue, and after six weeks the sponsor wants it opened to the whole department. Deflection is up, the testers are pleased, and the question on the table is whether to say yes.
I have sat in a lot of these meetings and I no longer treat the pilot result as the evidence. A pilot runs under conditions nobody admits are special. The maker reads every transcript each morning. The testers know it is a test and phrase their questions carefully. The knowledge source was tidied for the occasion. What I want to know is whether the workflow survives the week the maker is on leave and a new starter types something nobody anticipated.
So my readiness check ignores how good the answers were. It asks whether the workflow has owners, whether anyone can see it failing, and whether there is a date on which somebody must look at it again.
Name the four owners before anyone says autonomous
The first thing I ask for is four names. People, not teams, written into the design record. Who owns the outcome the agent is meant to produce. Who is the human escalation point when the agent hands off or gets it wrong. Who owns the data boundary, meaning what the agent may read and what it may write. And who reviews changes to instructions, topics, or enabled actions before they ship.
In most pilots these four roles collapse into the maker who built it. That is fine for a pilot. It stops being fine for a department, because the maker is rarely accountable for the case queue or the leave policy, and never the person who gets called when the agent quotes a customer the wrong refund window. When the sponsor says the team owns it, I write down the team lead's name and then check whether the team lead agrees.
The data boundary owner is the one most often missing. In Copilot Studio the maker chooses the knowledge sources and connector actions, and confirming which environment the agent lives in and what its Dataverse connection can actually reach takes a deliberate conversation with the platform team. In Agentforce the same question becomes which actions the agent can run and which objects the agent user can see. I once watched an agent answer an internal question with figures from a spreadsheet the asker should never have been able to open. Nobody was deciding the boundary, and that is a security question before it is an AI question.
Run the Friday afternoon test
The second check is a single scenario I ask the room to walk through out loud. It is ten past four on a Friday. The agent has just closed a case that was not resolved, or told an employee they can take leave they cannot, or attached a record to the wrong account. Who sees that first, and what can they do about it in the next thirty minutes?
The answers tell you almost everything. If the first person to see it is the customer, the workflow is not ready. If the maker checks the transcripts on Monday, the workflow is not ready. If they would switch it off, I ask who holds that permission and whether they are reachable on a Friday afternoon. Agents need a pause control that ordinary people can reach, and sometimes the honest answer is that nobody could suspend this one without deploying a change. That finding decides the meeting.
I also ask what happens to the work already in flight. An agent that ran for an hour before somebody stopped it has left records in several states, and the recovery owner needs a list of what it touched. In Agentforce that means somebody outside the build team can pull the session history for a given window. In Copilot Studio it means the escalation owner, not only the maker, can reach the transcripts and any flow run history. If that list cannot be produced on demand, recovery turns into an archaeology exercise.
Look for evidence the agent knows when to stop
A workflow is ready to grow when it hands off readily and the handoff lands somewhere. In the pilot transcripts I look for the conversations where the agent said it could not help, and count them. A pilot with zero handoffs is a warning rather than a success. Either the questions were too easy or the agent answered things it should have escalated.
Then I follow one handoff to the end. Where did it go, who picked it up, and what did they see? A handoff into a queue nobody watches is the same as no handoff. The receiver needs the conversation so far, the agent's stated reason for stopping, and a record showing a human took over, which is the substance of designing the handoff between human and agent work. If the receiving team learns about handoffs by scrolling a shared inbox, fix that before adding users.
The same question applies to approval steps. If the agent asks a person to approve before it writes a record, I read the pilot approvals. When every one was granted within seconds by the same person, the step has become a reflex click. Either give the approver something real to judge or remove the step and be honest about the autonomy already granted.
Put the review date in the change record
The last thing I ask for is a date. A calendar entry six to eight weeks after the wider rollout, with the four owners invited and an agenda already written. The agenda covers results against the intended outcome, the instructions and topics as they stand that day, and whether anyone used the escalation path and what they found.
The date matters because agent workflows drift. Someone edits the instructions to settle a complaint and nobody re-runs the earlier test questions. The knowledge source gains a new folder. An extra action gets enabled because a sales manager asked nicely. Six weeks after launch, production is running a different agent from the one that passed the pilot, and without a fixed date nobody notices until an incident review. The same discipline that goes into a release checklist for shared solutions applies here. The review date goes into the same change record as the rollout, with the name of the person who will chair it.
The check, in the order I ask it
This is the list I take into the meeting. It fits on one page and I go through it out loud with the sponsor and the maker in the room.
- The outcome owner, escalation owner, data boundary owner, and change reviewer are named people who have each confirmed they hold the role.
- The data boundary owner has read the agent's knowledge sources, connections, and enabled actions, and they match what the agent should reach.
- Someone other than the maker can suspend the agent within thirty minutes on a Friday, and has done it once in a test.
- The team can produce a list of everything the agent touched in a given time window.
- At least one pilot handoff has been traced end to end, and the receiver saw the conversation and the reason for the handoff.
- Approval steps in the pilot show evidence of a real decision, and any that were reflex clicks have been redesigned or removed.
- A review date is in the calendar with the four owners invited and the agenda written.
What to say when the answer is not yet
If every line holds, say yes to the wider rollout and put the date in. If two or more are missing, the answer I give is short. Come back in two weeks with the names filled in and the Friday test done. Sponsors take that better than you would expect, because it is a plan with a date rather than a refusal.
Before your next rollout meeting, send the Friday afternoon scenario to the sponsor and ask for a written answer. The reply usually tells you the decision before anyone sits down.



