Human in the Loop: Designing Approval That Is Not a Rubber Stamp

Human in the Loop: Designing Approval That Is Not a Rubber Stamp

Sixty proposals a day at 90 seconds each is 90 minutes of review nobody has. Four approval designs, what belongs on the screen, and how to read an approval rate.

See how Ward detects decisions that need a human, ranked by size

Get a demo → Take the 3-minute assessment
Contents

"Human in the loop" is usually a slogan, not a control

Every AI vendor says a human stays in the loop. In practice that often means a screen with a green button, shown to someone with 40 of them to clear and no way to check any of them.

That is not oversight. It is a signature collection process, and it is worse than full automation because it transfers accountability to a person who was never given the means to exercise it.

A real approval control has four properties: the approver can see the evidence, they can act on it in the time they have, rejecting is as easy as approving, and the rejection changes something. Most deployments have one of the four.

The rubber stamp is arithmetic, not attitude

Give a buyer 60 replenishment proposals a day. Each one needs 90 seconds of genuine review, which means checking the demand signal, the on-hand position, and the lead time. That is 90 minutes of concentrated work inside a job that already has a full week in it.

What happens is what always happens. By week three they approve in batches. By week six the approval rate is 98% and the control is decorative.

The fix is volume, not training. Cut proposals to what a person can actually review, rank by dollar impact, and auto-approve the bottom band under a hard cap. Twelve reviewed decisions beat 60 stamped ones, and the twelve are where the money is.

Four approval designs and when each fits

Approve each action. The default, and correct for high-value, low-volume, hard-to-reverse decisions. A markdown on a seasonal category. A vendor order above a threshold.

Approve the policy, monitor the actions. The human approves a rule and the agent applies it, with exceptions escalated. Correct for high-volume, low-value, reversible decisions. Replenishment on fast-moving staples belongs here.

Approve by exception. The agent acts unless a human intervenes within a window. Correct only where the action is reversible and the cost of delay exceeds the cost of a bad action. Rare in retail. Real in fulfillment.

Sample and audit. The agent acts, and a human reviews a random 5 to 10% after the fact. Correct for high-volume tier 1 output where the risk is a wrong report rather than a wrong transaction.

Most retailers apply design one to everything, which produces the rubber stamp, then conclude that human review does not work. The design was wrong for the volume, not the principle.

What has to be on the approval screen

Six things, and their absence is why approvers cannot actually check anything.

  • The recommendation, stated as an action with a quantity, not as an observation.
  • The evidence, including the query or the calculation, viewable in one click.
  • The size, in dollars or units, so the approver knows how much attention this one deserves.
  • The assumption that would have to be wrong for this to be a bad call.
  • What happens if you do nothing, which is the comparison the approver is actually making.
  • A reject reason field with three or four preset options, because reject reasons are the training data for everything that follows.

That last field is the one teams cut for scope and regret. Without reject reasons you know your approval rate and nothing about why.

See how Ward detects decisions that need a human, ranked by size

Get a demo →

Read the approval rate like a gauge

The approval rate is the cheapest quality metric in any agent deployment, and it reads in three bands.

Below 70%: the agent is wrong often enough that review is doing real work. Keep the human, fix the model or the data. Do not discuss automation.

70 to 90%: the healthy operating band. Reviewers are catching real errors and the agent is carrying real load.

Above 95%: two possible explanations, and they need different responses. Either the agent is genuinely reliable in this narrow band, in which case automate the band and keep review for the exceptions. Or nobody is reading, in which case your control does not exist.

Tell them apart by seeding. Insert a small number of deliberately bad proposals and measure the catch rate. If reviewers pass a proposal to order 40 weeks of supply of a discontinued SKU, you have your answer.

Design the escalation, not just the approval

An approval queue without escalation stalls. Proposals sit, the window closes, and the decision gets made by inaction, which is the one path nobody designed.

Set a time-to-decision for each proposal type, escalate to a named backup when it lapses, and make the default at expiry explicit. For orders the default is usually no action. For a markdown late in a season, no action is itself a costly decision and should escalate rather than expire.

Publish the expiry behavior. Approvers behave differently when they know that not clicking is a choice with a consequence.

Why the CIO designs this and the business staffs it

Approval design determines whether your audit trail means anything. A logged approval from a reviewer who could not see the evidence is a record of a process that did not happen, and that is worse for you in a review than no record at all.

IT owns the mechanics: the queue, the evidence display, the logging, the escalation, the identity of the approver. The business owns the thresholds and staffs the reviewers. Both own the approval rate, and it should be on a monthly report that someone reads.

The question to ask any vendor: show me the approval screen, tell me your median review time across your customers, and tell me your override rate. A vendor who does not measure override rate is not measuring whether their oversight design works.

How Ward handles the loop

Ward files cases, not proposals to stamp. Each card carries the finding, the query, the dollar size, and the action, ranked so the top three are the three worth the time. Volume is capped per recipient because a card nobody reads is worse than no card.

Cases close when the KPI moves, not when someone clicks. Lane assist, not autopilot, and not a green button either.

See how Ward detects decisions that need a human, ranked by size

Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.

human in the loop approval workflow AI agents controls oversight

Questions about decisions that need a human, ranked by size.

Four properties. The approver can see the evidence, they can act in the time they have, rejecting is as easy as approving, and rejection changes something downstream. Most deployments have one of the four, which produces a screen with a green button shown to someone with 40 of them to clear. That is signature collection, not oversight.

Arithmetic, not attitude. Sixty proposals a day at 90 seconds of genuine review each is 90 minutes of concentrated work inside a full job. By week three approvals happen in batches and by week six the approval rate is 98%. The fix is volume: rank by dollar impact, send what a person can review, and auto-approve the bottom band under a hard cap.

Approve each action, for high-value low-volume irreversible decisions. Approve the policy and monitor the actions, for high-volume low-value reversible ones like staple replenishment. Approve by exception, where the agent acts unless someone intervenes in a window, which is rare in retail. And sample-and-audit, reviewing a random 5 to 10% after the fact for read-only output.

Six things: the recommendation stated as an action with a quantity, the evidence including the query in one click, the size in dollars or units, the assumption that would have to be wrong, what happens if you do nothing, and a reject reason field with preset options. The reject reasons get cut for scope and are the only thing that tells you why the approval rate is what it is.

Below 70%, the agent is wrong often enough that review is doing real work, so fix the model or the data and do not discuss automation. Between 70 and 90% is the healthy band. Above 95% means either genuine reliability in a narrow band, which you can automate, or that nobody is reading. Tell them apart by seeding deliberately bad proposals and measuring the catch rate.

From the article to the product.

How this topic maps to what Ward does, who it’s for, and the alternatives buyers benchmark against.

Your stores are generating data right now.

Ward turns it into decisions. First insight cards in 48 hours.

デモを予約

御社のデータに何が隠れているか、確認する。

お客様のオペレーションについてお聞かせください。デモまたはPoCをご提案します。

ステップ 1/3
解決したい課題は?
ステップ 2/3
御社について
ステップ 3/3
ご連絡先