One clean cycle is the weakest evidence you can bring

A parallel run that matches to the penny on its first attempt tells you one thing. The new system calculated correctly for the people who happened to be paid that period. Most periods are quiet. The cases that break a payroll engine are starters and leavers, retroactive adjustments, workers with two assignments and anyone sitting on a statutory threshold, and none of those are guaranteed to turn up in the month you picked.

Three consecutive cycles is the honest floor for a monthly payroll. Cycle one finds the structural problems, the wrong earning codes and mapping gaps that hit dozens of employees at once. Cycle two tests whether your fixes held and brings in the cases you built deliberately. Cycle three proves stability, because a configuration that was right twice under different conditions is harder to argue with.

Weekly and fortnightly payrolls need more periods to cover the same ground. Three weekly cycles miss a month end, most retro scenarios and any quarterly statutory event. Run enough consecutive periods to cover a full calendar month plus one period containing a known statutory boundary. If your third parallel was the first clean one, you have one clean run and two failures, and the pack should say so.

Build the hard cases in rather than hoping they show up

Run the full population if the tooling allows it, then check which case types are actually in it. A parallel over a random sample of 200 employees will match beautifully and prove very little. The value of the run comes from whether it contains the situations that exercise the rules you least trust.

Six case types earn a deliberate place. A starter who joined mid-period, so pro-rating on a first payslip gets tested. A leaver who left mid-period, with final pay, notice and untaken leave paid out. Someone with retroactive pay crossing a period boundary, because retro is where engines diverge most and legacy behaviour is least documented. Someone holding two concurrent assignments, or a worker whose cost centre changed part way through the period. Anyone whose statutory deductions land on a threshold boundary, such as a pension auto-enrolment trigger or a social security band edge. And any population paid on a different frequency, because frequency-specific annualisation rules go untested until the first live run.

Where a case type genuinely does not exist in the population for the period you are running, build a synthetic employee that exercises it and mark the record so it can never reach a payment file. That is the same discipline behind a full migration rehearsal. A rehearsal that only touches easy data confirms the process starts and finishes, and says nothing about the rules that worry you.

Variance triage is the actual work

Running the parallel takes a day. Reconciling it takes weeks, and that is the part project plans keep underestimating. Decide your variance classes before you see a single number, because a tolerance agreed afterwards lands wherever the deadline needs it.

Set a rounding tolerance per pay component per employee rather than on the total, because aggregate matching hides offsetting errors where an overpayment on one code cancels an underpayment on another. Two pence a component is defensible for most engines. Whatever figure you pick, write it down, have the payroll manager and the finance controller sign it before cycle one, and hold to it. Anything outside tolerance is a defect until somebody proves otherwise, and proving otherwise means a rule reference or a written decision. Payroll cutover reconciliation needs the same per-component detail.

The pressure to explain variances away arrives around week three, when the steering group asks whether the date still holds. Somebody will propose that a cluster of small differences is a legacy quirk and can be accepted in bulk. Make them itemise it. An accepted difference needs a named approver, a reason and an expected value, recorded in the same register as the defects. Bulk acceptance is how a systematic error reaches production with nobody's name against it.

The 212 pounds that turned out to be a pension rule

On a Workday HCM parallel at a facilities company with about 4,200 employees, our second cycle returned net pay matching for everyone except 37 people. Thirty-one sat under five pence, inside the agreed tolerance. Five were pension differences of a few pounds each, all in the same direction. One was 212 pounds.

The instinct in the room was to close the 212 pounds as legacy error, because the record was unusual and the deadline was nine days out. We opened it anyway. She held two concurrent assignments, one at 0.6 FTE in a regional office and one at 0.4 FTE covering a contract site. Legacy assessed pension on her combined earnings. Our configuration applied the qualifying earnings threshold per assignment, so each job fell below the trigger and she was under-deducted on both.

She was the only person in that period with two live assignments. Across the year the dual-assignment population was closer to ninety, growing every autumn when sites were staffed from people already on the books. Had she sat outside our scope, we would have gone live under-deducting pension for ninety people, surfacing only when somebody next reconciled provider contributions, which at that company happened quarterly. A better scope question asks which employee exercises each rule you are unsure about.

Who signs and what they need in front of them

Sign-off belongs with the payroll manager for calculation accuracy, the finance controller for cost and posting, and the HR systems lead for the data that fed the run. The programme manager does not sign, because their incentive is the date and the date is the pressure this process exists to resist.

Those three need a variance register where every line carries a disposition, an accepted-difference list with a named approver and reason against each entry, and a coverage table naming the employee number that exercised each promised scenario. Add the tolerance definition, signed and dated before cycle one ran. That pack is what an auditor asks for when somebody queries a deduction eighteen months later.

A slide reading 99.8 percent match is not sign-off evidence. Ask whoever presents it which employees make up the remaining 0.2 percent and what was decided about each. If nobody can answer, the parallel has not finished, whatever the calendar says. Configuration drift in a Workday comp cycle fails the same way, behind a summary number nobody opened.

What a parallel does not prove

A clean parallel proves calculation fidelity against the old engine for the periods tested. It does not prove the old engine was right. If legacy has applied a salary sacrifice rule incorrectly since 2019, a perfect match reproduces that error and hands you a signed document saying so. Any rule already under review should be tested against the policy, not the previous output.

The run also says nothing about what happens downstream. General ledger postings into finance, the contribution file to the pension provider, the payment file to the bank, statutory filings, payslip rendering and the reporting layer all consume payroll results through code paths the parallel never touches. Those need their own cycles and owners, with the integration side given the scrutiny described in integration that survives a reorg.

And no parallel reproduces the first live run. Yours ran with a week of slack and the freedom to stop and investigate. The live run has a Tuesday cut-off, a bank submission deadline, a payroll officer covering two other tasks and no option to pause. Rehearse the timings against the clock, with the people who will be at the keyboard.

Before your next parallel opens, take the rules you are least confident about and write an employee number beside each one. Where you cannot, that rule is not being tested, and the run will come back clean for reasons unrelated to whether the configuration is correct.