Human in the Loop: Designing Approval That Is Not a Rubber Stamp
Sixty proposals a day at 90 seconds each is 90 minutes of review nobody has. Four approval designs, what belongs on the screen, and how to read an approval rate.
See how Ward detects decisions that need a human, ranked by size
Get a demo → Take the 3-minute assessmentContents
- "Human in the loop" is usually a slogan, not a control
- The rubber stamp is arithmetic, not attitude
- Four approval designs and when each fits
- What has to be on the approval screen
- Read the approval rate like a gauge
- Design the escalation, not just the approval
- Why the CIO designs this and the business staffs it
- How Ward handles the loop
"Human in the loop" is usually a slogan, not a control
Every AI vendor says a human stays in the loop. In practice that often means a screen with a green button, shown to someone with 40 of them to clear and no way to check any of them.
That is not oversight. It is a signature collection process, and it is worse than full automation because it transfers accountability to a person who was never given the means to exercise it.
A real approval control has four properties: the approver can see the evidence, they can act on it in the time they have, rejecting is as easy as approving, and the rejection changes something. Most deployments have one of the four.
The rubber stamp is arithmetic, not attitude
Give a buyer 60 replenishment proposals a day. Each one needs 90 seconds of genuine review, which means checking the demand signal, the on-hand position, and the lead time. That is 90 minutes of concentrated work inside a job that already has a full week in it.
What happens is what always happens. By week three they approve in batches. By week six the approval rate is 98% and the control is decorative.
The fix is volume, not training. Cut proposals to what a person can actually review, rank by dollar impact, and auto-approve the bottom band under a hard cap. Twelve reviewed decisions beat 60 stamped ones, and the twelve are where the money is.
Four approval designs and when each fits
Approve each action. The default, and correct for high-value, low-volume, hard-to-reverse decisions. A markdown on a seasonal category. A vendor order above a threshold.
Approve the policy, monitor the actions. The human approves a rule and the agent applies it, with exceptions escalated. Correct for high-volume, low-value, reversible decisions. Replenishment on fast-moving staples belongs here.
Approve by exception. The agent acts unless a human intervenes within a window. Correct only where the action is reversible and the cost of delay exceeds the cost of a bad action. Rare in retail. Real in fulfillment.
Sample and audit. The agent acts, and a human reviews a random 5 to 10% after the fact. Correct for high-volume tier 1 output where the risk is a wrong report rather than a wrong transaction.
Most retailers apply design one to everything, which produces the rubber stamp, then conclude that human review does not work. The design was wrong for the volume, not the principle.
What has to be on the approval screen
Six things, and their absence is why approvers cannot actually check anything.
- The recommendation, stated as an action with a quantity, not as an observation.
- The evidence, including the query or the calculation, viewable in one click.
- The size, in dollars or units, so the approver knows how much attention this one deserves.
- The assumption that would have to be wrong for this to be a bad call.
- What happens if you do nothing, which is the comparison the approver is actually making.
- A reject reason field with three or four preset options, because reject reasons are the training data for everything that follows.
That last field is the one teams cut for scope and regret. Without reject reasons you know your approval rate and nothing about why.
See how Ward detects decisions that need a human, ranked by size
Get a demo →Read the approval rate like a gauge
The approval rate is the cheapest quality metric in any agent deployment, and it reads in three bands.
Below 70%: the agent is wrong often enough that review is doing real work. Keep the human, fix the model or the data. Do not discuss automation.
70 to 90%: the healthy operating band. Reviewers are catching real errors and the agent is carrying real load.
Above 95%: two possible explanations, and they need different responses. Either the agent is genuinely reliable in this narrow band, in which case automate the band and keep review for the exceptions. Or nobody is reading, in which case your control does not exist.
Tell them apart by seeding. Insert a small number of deliberately bad proposals and measure the catch rate. If reviewers pass a proposal to order 40 weeks of supply of a discontinued SKU, you have your answer.
Design the escalation, not just the approval
An approval queue without escalation stalls. Proposals sit, the window closes, and the decision gets made by inaction, which is the one path nobody designed.
Set a time-to-decision for each proposal type, escalate to a named backup when it lapses, and make the default at expiry explicit. For orders the default is usually no action. For a markdown late in a season, no action is itself a costly decision and should escalate rather than expire.
Publish the expiry behavior. Approvers behave differently when they know that not clicking is a choice with a consequence.
Why the CIO designs this and the business staffs it
Approval design determines whether your audit trail means anything. A logged approval from a reviewer who could not see the evidence is a record of a process that did not happen, and that is worse for you in a review than no record at all.
IT owns the mechanics: the queue, the evidence display, the logging, the escalation, the identity of the approver. The business owns the thresholds and staffs the reviewers. Both own the approval rate, and it should be on a monthly report that someone reads.
The question to ask any vendor: show me the approval screen, tell me your median review time across your customers, and tell me your override rate. A vendor who does not measure override rate is not measuring whether their oversight design works.
How Ward handles the loop
Ward files cases, not proposals to stamp. Each card carries the finding, the query, the dollar size, and the action, ranked so the top three are the three worth the time. Volume is capped per recipient because a card nobody reads is worse than no card.
Cases close when the KPI moves, not when someone clicks. Lane assist, not autopilot, and not a green button either.
See how Ward detects decisions that need a human, ranked by size
Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.