Every agent description reads the same
Every finance and operations vendor now describes agents that reason, decide and act on a user's behalf. Cover the logos on four of those descriptions and most architects cannot tell which suite is which. The verbs match, the diagrams match, the claimed outcome matches. A buyer has nothing left to compare, so the comparison has to be built by testing the claim rather than reading it.
Oracle shops sit in a specific version of this. Agent naming inside Fusion moves faster than anybody's documentation, ours included, and what appears in your pods depends on your release level and your entitlement. So this piece names no features and no dates. Read your own release notes and subscription to establish what you actually have.
Four questions cover the agent itself and three cover the finance context it would run in. None of them need a product name, which is why they survive the next round of renaming. If you have already decided to switch something on, the rollout mechanics live in a separate guide.
Start with what the agent is allowed to touch
The first question is what the agent acts on. An agent that drafts a journal for a person to review carries one kind of risk. An agent that posts the journal carries another. Both get described in identical language, and the second one changes your financial records with nobody in the path.
Draft and suggest cost less to be wrong about. A bad draft gets rejected by whoever reviews it and the damage stops there. A bad posting travels into subledger balances, into a report someone presents on Friday, and into a period already signed off. Settle which of the two you are buying before anyone talks about price.
Ask the vendor to name every write the agent can perform, by object, not by capability statement. If the answer stays at the level of a capability statement, keep asking. A team who built the thing can tell you which records change. A team carrying a deck usually cannot.
Whose access it acts under decides most of it
The second question separates deployable from not deployable. An agent executing under the access of the named person who asked for the work leaves an auditable record. That person's identity sits on the action, their permissions bound what the agent could reach, and your access reviews already cover them. An agent executing under one broad service identity shared across the tenant hands you an action with no human attached.
That is not a matter of taste. If the agent runs under a general-purpose integration account, your auditor has no route from a posting back to a person, and your quarterly access review has nothing to look at. We have watched evaluations end on that single answer, correctly. Service identity varies across platforms and finance is the least forgiving case.
Ask which identity lands in the audit record, and whether it can be constrained by the role assignments you already maintain. Then ask what happens when the requesting person lacks a permission the agent needs. Some implementations stop and raise it, which is right. Some quietly escalate to a broader identity, which removes the control you thought you had bought.
What it leaves behind for somebody to read later
The third question is evidence. Six weeks later, can you reconstruct what the agent did and what it decided on? A log line saying an agent ran does not get you there. You want the inputs it read, the reasoning it applied, the option it passed over, and the person or schedule that set it going.
Most ERP audit records were designed to capture who changed a field and when. That is thinner than an agent decision needs, because the part questioned later is the basis rather than the change. If the vendor's answer is that the platform writes its standard audit rows, you will be reconstructing intent from memory in front of somebody paid to doubt you.
Pick one action the agent would take and ask the vendor to show the full evidence trail for it in a running system. Not a screenshot of a monitoring page. The actual record a reviewer would pull, exportable if possible, retained as long as the rest of your financial evidence.
The behaviour you want when it is unsure
The fourth question is what the agent does with uncertainty. The useful behaviour is stopping and asking a person. The dangerous behaviour is proceeding with a confident answer that happens to be wrong, because a confident wrong posting is indistinguishable from a correct one until somebody reconciles.
Ask how the agent expresses doubt and what threshold sends work back to a human. Then ask what happens when it is wrong and confident at once. Every system has that case somewhere. A vendor who says their agent does not is telling you they have not measured it.
The answer you want names a specific stop, a specific queue and a role who receives the item. The answer to distrust describes confidence scoring with no route to a person. Pair it with who can switch the whole thing off inside a minute, because an agent with no pause is a design gap rather than an operating model.
Close has a deadline and duties are a control
Everything above applies to agents anywhere. Three things make finance different. Period close attaches a deadline to errors, so a mistake caught on day two costs an adjusting entry and an irritated accountant, while the same mistake found after close costs a reopened ledger, a restated period, a conversation with your auditor and sometimes a filing you have to correct. Ask where in the calendar the agent may act, and whether you can freeze it for the last days of close.
Segregation of duties gets collapsed quietly. Your organisation spent real money separating the person who creates a supplier from the person who pays one. An agent granted the combined access of several roles so it can finish a task end to end defeats that separation, and design documents rarely say so out loud. List the roles the agent needs and run them through the same conflict matrix you use for people. Duty conflicts are hard enough with one person wearing four hats, and an agent can hold more at once.
Statutory and audit obligations outlast the feature. An external party may eventually ask how one posting was determined, and the answer has to hold years later, possibly after the agent has been reconfigured or removed. Retention of the decision evidence matters as much as the evidence existing on the day.
Measure the work before you change it
Now the business case. Every agentic claim arrives with effort saved attached, and effort saved cannot be falsified without a measurement taken beforehand. If you do not know how many hours your team spends matching supplier invoices today, you cannot show afterwards whether the agent helped, and the review meeting turns into an argument about a number nobody can reproduce twice.
Almost nobody takes that baseline, which is why it is the highest-value preparation available to you. It costs a few weeks of deliberately boring measurement before you enable anything. Count the transactions, time the steps, count the exceptions and the rework, and note who does them today. If a full baseline is more than a first pass can carry, take the shorter version from an agent readiness check and measure one process instead of all of them.
So here is the question to put to any vendor making agentic claims about a finance suite. Ask them to walk you through one transaction their agent completed in a real tenant, from the trigger to the posted record, naming the identity in the audit row, the evidence a reviewer could pull six months later, and the point at which it would have stopped and asked a person. Vendors who have built it can do that in ten minutes. Vendors who have not will offer a roadmap session, and that answer tells you something too.


