The records that updated on one side only
A purchase approval flow ran on a Tuesday night, stamped forty-one requests as approved in Dataverse, and posted none of them into the finance system. The first finance call failed. Someone had set the next action to run after that failure while testing, months earlier, and nothing downstream checked the result, so the loop carried on and the run finished reporting success.
Nine days passed before anyone looked. Finance found it during month-end reconciliation. The approvers had moved on, the requesters believed their orders were placed, and two suppliers were waiting on purchase orders that existed in one system only. Untangling it cost two people most of a day.
A flow that fails outright is the cheap outcome. The run goes red, the run history names the action that broke, and the data is either untouched or obviously unfinished. A flow that half-succeeds leaves two systems disagreeing and no signal at all, and the cost lands on whoever trips over the inconsistency weeks later.
Configure run after is the panel most makers never open
Every action in a cloud flow carries settings that say when it should run, and by default an action runs only if the one before it succeeded. In the Power Automate designer you reach them from the menu on the action itself, under Configure run after, where the conditions on offer are success, failure, timeout and skip.
Ticking failure alone is the mistake I see most. An action that never returns a response times out rather than failing, and an action behind one that already failed gets skipped rather than run. A handler wired to failure only sits out both cases, and the flow ends without the step meant to clean up.
Catching a failure also changes what the run reports. Once a handler takes the failed path and nothing after it fails, the run can finish with a status of success, which is what you want for a problem the flow recovered from and what hides one it did not. Decide which of the two you have before you write the handler.
Scope gives you one handler instead of one per action. Put the actions that belong together inside a Scope, and the Scope reports a single status for the group, so a handler placed after it with run after set to failure and timeout answers for anything inside. It also gives you somewhere to record which group of work did not finish.
Retry the network, not the business rejection
Many connector actions expose a retry policy in their settings, and the platform repeats the call on certain classes of failure before reporting one. The counts, the intervals and which failures qualify vary by action, connector and release, so read the setting in front of you and confirm current behaviour in the documentation and in your own tenant.
The distinction that earns its keep is transient against final. A dropped connection or a throttled response has a fair chance of succeeding second time round, and a retry saves you an incident. A validation error or a missing permission fails the same way every time, so retrying it spends run duration for nothing and risks duplicate work if the first call landed before its response went missing.
When a run reaches a state you would not want continued, end it on purpose. The Terminate action stops the flow with a status you choose, and the failed option carries a code and a message you write yourself. Name the record and the reason in that message, because it is what the person reading run history has to work from.
Ask what makes an action safe to run twice
A retried flow that creates a second record costs more than the failure it was retrying. Every action that creates something needs an answer to one question, which is what makes it safe to run twice. Ask it of each create, each send and each post before the flow goes near production.
The practical answer is a key the receiving system enforces. Upsert against a business key rather than creating blind, and let the target reject the duplicate instead of asking flow logic to notice it. Reading first and creating second leaves a gap that a retry lands in. Durable business keys make the upsert possible, and choosing one is a data decision.
Some actions cannot be made safe. An email goes out once, a payment posts once, a supplier portal accepts the submission once. For those, write a state onto a record you control before you fire the action and check that state on the way in, so a second run can tell the work already happened.
Decide what step four of seven leaves behind
For any flow with more than a couple of writes, write down what the world looks like if the run stops after each step. It takes twenty minutes and it is the most useful twenty minutes a maker spends. The answer is usually that two systems disagree, and usually nobody had thought it through.
Order the actions so the reversible work happens first and the irreversible work happens last, wherever the process allows. Validate, write the records you can correct, then send the notification or post the transaction. When the irreversible step has to come first because everything downstream needs an identifier from it, you owe the flow a compensating action.
A failure inside a loop needs a decision rather than a default. Stop the batch on the first failure, or record the failure and keep going. Continuing without noticing is how the purchase order flow stayed quiet for nine days. Collect the failed items, then terminate with a failed status if that collection has anything in it, so run history says the batch was partial. Automation owners who lose the thread usually lost it in a loop.
The failure notice goes to whoever left
Failure notices default to the flow owner, so a flow built by a contractor who finished in March is a flow whose failures arrive in a mailbox nobody opens. Check who receives them before you check anything else, because flow ownership debt accumulates quietly and surfaces on the morning the flow matters.
For anything that moves money or touches a customer, email is the wrong destination. Route the failure into a queue or raise a ticket with an assignment group, so there is a record with a state and someone whose job is clearing it. Same argument as error handling that pages someone, and it holds inside a cloud flow too.
Run history is retained for a limited window that depends on the platform and your licence, so confirm the current figure rather than assuming one. Diagnosis months later depends on what got written at the time. The inputs and outputs of an expired run are gone, and what remains is the row you inserted and the status you stamped on the record.
Ownership is the last piece. A flow with no named owner in a register is an incident waiting for a date, and a flow running under a personal account dies with that account, connections and all. Service accounts across platforms need their own handling. Before a flow takes on production work, this is the list I walk with the maker.
- Every action that writes to another system has a handler for failure, timeout and skip.
- Every path that leaves work unfinished ends in a Terminate with a written message.
- Every create is an upsert on a business key or gets rejected by a constraint in the target.
- The irreversible action runs last, or has a named compensating step.
- A failure inside a loop either stops the batch or lands somewhere a person reads.
- Failure notices reach a queue or a ticket rather than one person's mailbox.
- The flow runs under a service account and has a named human owner written down.
Open the flow that touches money first
Find the flow in your tenant with the highest run count against a finance or customer system. Open every action in it that writes somewhere and look at what runs after that action fails. Then ask the maker what state the world is in if the run stops halfway, and wait for the answer.
Most of the time the fix is an hour of run after settings, one Scope and one Terminate. That hour is available now. It stops being available on the morning finance asks why the ledger and the approvals disagree, and by then somebody else is deciding what you work on.


