Skip to content
Webb Technologies

AI integration · Human review

Designing human review and approval into AI workflows

"A person will review it" is the most common answer to the question of what happens when an AI workflow gets something wrong. It's the right answer, but only if the review is designed. A queue of AI suggestions that people click through without reading adds cost and a false sense of safety. This guide covers how to decide what needs review, how to route the right items to people, what a review screen needs, how approval gates for actions work, and how to reduce review over time without losing control.

Webb TechnologiesUpdated 9 min read

Drafted with AI assistance; facts checked against any sources cited. General guidance, not advice for your situation; verify before relying on it.

On this page

In short

  • Match the review pattern to the consequence of an error: sampled checks for internal suggestions, approval before anything leaves the company or touches money.
  • Route items to people on signals you can verify, such as failed validation and business rules, rather than on the model's own confidence alone.
  • Make reviewing faster than doing the work by hand, or reviewers will either skip it or rubber-stamp it.
  • Log every decision and correction, keep a random sample under review indefinitely, and let categories graduate to less review only on measured results.

Why review steps quietly stop working

Review steps rarely fail in a way anyone notices. They erode. The usual causes are predictable:

  • Rubber-stamping. When the AI is right most of the time, people start approving without reading. The occasional error sails through with a human signature on it.
  • Review slower than the work. If checking a suggestion takes as long as doing the task by hand, reviewers go back to doing it by hand, or stop checking.
  • No context on screen. A reviewer who can't see the source document next to the output can't judge it, so they guess or trust.
  • Everything or nothing. Every item goes to a person forever, so the workflow never saves time; or review is dropped entirely after a good first month.
  • Corrections thrown away. Reviewers fix the same kind of mistake every day and nothing improves, because the fixes are never captured.

Each of these is a design problem, and each is cheaper to address before launch than after.

Decide what needs review, by consequence

Start by listing every output the workflow produces and what happens if it's wrong. The review pattern follows from that, not from how confident anyone feels about the model.

Matching the review pattern to the consequence of an error

Internal suggestion (a category, a tag, a summary for a colleague)

If it's wrong
A person notices and corrects it in passing
Review pattern
Sampled review after the fact

Fields written to a system of record

If it's wrong
Bad data spreads to reports and downstream systems
Review pattern
Review until accuracy is measured, then review exceptions and a sample

Drafts that go outside the company

If it's wrong
A customer or supplier receives something wrong
Review pattern
Approval of each item before it's sent

Actions involving money, master data or anything safety-related

If it's wrong
Financial, compliance or safety impact
Review pattern
Approval by a named person with the authority; two people where your policies already require it

Four review patterns

Common human review patterns in AI workflows

Approve before it happens

How it works
Nothing takes effect until a person approves it
Good for
Outbound messages, payments, master-data changes, early days of any workflow

Exception review

How it works
Items that fail checks or rules go to a person; the rest go straight through
Good for
High-volume extraction and classification once accuracy is measured

Sampled review

How it works
A random share of straight-through items is checked after the fact
Good for
Catching drift and silent errors the checks miss

Two-person approval

How it works
One person prepares or accepts, a second approves
Good for
Anything your existing policies already require two signatures for

Most workflows combine them. A typical progression is approval of everything at launch, then exception review plus sampling for the categories that measure well, with approval gates kept permanently for outbound and hard-to-reverse actions.

Route on signals you can trust

Exception review depends on deciding which items a person should see. It's tempting to ask the model how confident it is and route anything below a threshold. A model's own statement of confidence is not a reliable measure of whether it's right (research on this has found models tend to be overconfident when they state it), so use it, if at all, alongside signals your code can verify.

  • Validation failures. A required field is missing, a part number isn't in the item master, a date is impossible, line totals don't add up.
  • Business rules. Anything over a value threshold, a new supplier or customer, a first order for a part, a request that mentions a complaint or a legal term.
  • Measured accuracy by category. Categories or fields that have scored poorly on your evaluation set always go to a person until they improve.
  • Disagreement. Two extraction passes, or two models, that return different values for the same field.
  • Unfamiliar input. A document layout, sender or language the workflow hasn't handled before.

Record why each item was routed to review and show that reason to the reviewer. "Part number not found in item master" focuses attention far better than a generic "needs review".

A reference design

ai-workflow / review and approvaldiagram
The AI step produces a structured result. Code checks it against the schema, master data and business rules and decides the route: straight through, to a review queue with the reason attached, or to an approval gate for actions. Reviewers and approvers work from a screen that shows the source alongside the result. Only validated or approved results reach the system of record, and every decision is logged.

Routing and approval are enforced in code; the model can't skip a queue by what it writes.

A review screen people will actually use

The review screen decides whether review is real. Treat it as a product in its own right, and test it with the people who'll use it.

Review screen essentials

  • The source document or message side by side with the AI's output, with the relevant passage highlighted where possible.
  • The reason the item was routed to review, stated plainly.
  • Edit a single field in place, rather than rejecting and re-keying the whole record.
  • A short list of rejection reasons, so rejections can be counted and grouped.
  • Keyboard shortcuts for accept, edit and next, for high-volume queues.
  • No "approve all" for items behind an approval gate, and no pre-ticked approvals.
  • For actions, exactly what will happen on approval: the recipient, the record, the values before and after.

Watch for signs of rubber-stamping: approvals that take a second or two on items that need reading, or a reviewer who never edits anything. A simple countermeasure is to occasionally include a known-wrong item from your test set in the queue and check that it gets caught. Treat this as a check on the design, not on the people.

Approval gates for actions

When a workflow or an AI agent can do something, such as sending an email, creating an order or updating a customer record, the approval has to be enforced by your code, not requested in the prompt.

  • Pause and queue. A gated action is stored with its exact payload and waits. Nothing runs until an approval is recorded.
  • Approver authority from existing roles. Map who can approve what to the roles or groups you already manage in your identity provider, not to a separate list inside the AI tool.
  • Specific and time-limited. An approval covers one action with one payload. If the payload changes, it needs approving again; if it sits too long, it expires.
  • Re-validate at execution. Check the payload again when the approved action runs, in case the underlying record changed in the meantime.
  • Separate requester and approver where money or policy is involved, as you would for the same action done by hand.
  • Never approve on timeout. If nobody approves, the action waits or expires. It doesn't go ahead by default.

Our guide to putting AI agents to work on internal processes covers tool permissions and approval levels for agents in more detail, and securing internal apps with single sign-on and least privilege covers the role design approvals rely on.

Capture every decision and measure it

Every review is a labeled example. Capturing it is what lets the workflow improve and lets you reduce review safely.

What to log for each reviewed item

  • The input, the AI's output, and the model and prompt version that produced it.
  • Why the item was routed to review.
  • The reviewer's decision, every edited field with before and after values, and the rejection reason if any.
  • Who reviewed it, when, and how long it took.

Review measures worth tracking weekly

Accepted without edits

What it tells you
How often the AI's output is right as it stands, per category or field

Edit rate per field

What it tells you
Which fields need a better prompt, extraction step or validation rule

Review time per item

What it tells you
Whether review is actually faster than the manual process

Errors found in sampled items

What it tells you
Whether straight-through processing is as accurate as assumed

Queue age

What it tells you
Whether review is keeping up, or items are waiting too long

Add corrected items to your evaluation set, so the next prompt or model change is tested against the mistakes your reviewers actually found. Our guide to adding AI to document and email workflows covers that evaluation loop.

Reducing review without losing control

  1. Agree in advance what a category or field must achieve, over how many items, before it moves from full review to exception review. Write it down with the process owner.
  2. Move one category at a time, and keep sampling it at a meaningful rate afterward.
  3. Return a category to full review automatically if sampled errors or edit rates rise past an agreed level.
  4. Put everything back under full review for a period after any change to the model, prompt or extraction step, and graduate it again on the new results.
  5. Keep approval gates on outbound, financial and master-data actions permanently, whatever the accuracy.

Adding review to an AI workflow?

Tell us what the workflow produces and who checks it today. On a scoping call we'll talk through which outputs need review, where approval gates belong, what the review screen should show, and what a fixed-price build would cover.

Request a scoping call

Not ready to talk yet? Draft a project brief first 

A one-page review design checklist

Agree these before launch

  • Every output the workflow produces, what happens if it's wrong, and the review pattern for each.
  • The signals that route an item to review, and the reason shown to the reviewer.
  • Which actions sit behind an approval gate, and which roles can approve them.
  • A review screen tested with the people who'll use it.
  • What gets logged for each decision, and where it's stored.
  • The measured results a category needs before its review is reduced, and the sampling rate after.
  • The owner of the review queue and what happens when it backs up.

Our experience includes document and email processing, classification, summarization, automated workflows and AI agents, with integrations across the Anthropic Claude, OpenAI, Google Gemini and xAI Grok APIs. We design human review and approval steps into any workflow where output affects customers or money. See AI integration for how we'd approach your workflow, and compare your options if you're deciding whether to build it in-house or with outside help.

Sources

Want to talk through your version of this?

A 30-minute call. You leave with a clear approach and the real risks, whether or not you hire us.