On this page
In short
- Match the review pattern to the consequence of an error: sampled checks for internal suggestions, approval before anything leaves the company or touches money.
- Route items to people on signals you can verify, such as failed validation and business rules, rather than on the model's own confidence alone.
- Make reviewing faster than doing the work by hand, or reviewers will either skip it or rubber-stamp it.
- Log every decision and correction, keep a random sample under review indefinitely, and let categories graduate to less review only on measured results.
Why review steps quietly stop working
Review steps rarely fail in a way anyone notices. They erode. The usual causes are predictable:
- Rubber-stamping. When the AI is right most of the time, people start approving without reading. The occasional error sails through with a human signature on it.
- Review slower than the work. If checking a suggestion takes as long as doing the task by hand, reviewers go back to doing it by hand, or stop checking.
- No context on screen. A reviewer who can't see the source document next to the output can't judge it, so they guess or trust.
- Everything or nothing. Every item goes to a person forever, so the workflow never saves time; or review is dropped entirely after a good first month.
- Corrections thrown away. Reviewers fix the same kind of mistake every day and nothing improves, because the fixes are never captured.
Each of these is a design problem, and each is cheaper to address before launch than after.
Decide what needs review, by consequence
Start by listing every output the workflow produces and what happens if it's wrong. The review pattern follows from that, not from how confident anyone feels about the model.
Matching the review pattern to the consequence of an error
Internal suggestion (a category, a tag, a summary for a colleague)
- If it's wrong
- A person notices and corrects it in passing
- Review pattern
- Sampled review after the fact
Fields written to a system of record
- If it's wrong
- Bad data spreads to reports and downstream systems
- Review pattern
- Review until accuracy is measured, then review exceptions and a sample
Drafts that go outside the company
- If it's wrong
- A customer or supplier receives something wrong
- Review pattern
- Approval of each item before it's sent
Actions involving money, master data or anything safety-related
- If it's wrong
- Financial, compliance or safety impact
- Review pattern
- Approval by a named person with the authority; two people where your policies already require it
| Output | If it's wrong | Review pattern |
|---|---|---|
| Internal suggestion (a category, a tag, a summary for a colleague) | A person notices and corrects it in passing | Sampled review after the fact |
| Fields written to a system of record | Bad data spreads to reports and downstream systems | Review until accuracy is measured, then review exceptions and a sample |
| Drafts that go outside the company | A customer or supplier receives something wrong | Approval of each item before it's sent |
| Actions involving money, master data or anything safety-related | Financial, compliance or safety impact | Approval by a named person with the authority; two people where your policies already require it |
Four review patterns
Common human review patterns in AI workflows
Approve before it happens
- How it works
- Nothing takes effect until a person approves it
- Good for
- Outbound messages, payments, master-data changes, early days of any workflow
Exception review
- How it works
- Items that fail checks or rules go to a person; the rest go straight through
- Good for
- High-volume extraction and classification once accuracy is measured
Sampled review
- How it works
- A random share of straight-through items is checked after the fact
- Good for
- Catching drift and silent errors the checks miss
Two-person approval
- How it works
- One person prepares or accepts, a second approves
- Good for
- Anything your existing policies already require two signatures for
| Pattern | How it works | Good for |
|---|---|---|
| Approve before it happens | Nothing takes effect until a person approves it | Outbound messages, payments, master-data changes, early days of any workflow |
| Exception review | Items that fail checks or rules go to a person; the rest go straight through | High-volume extraction and classification once accuracy is measured |
| Sampled review | A random share of straight-through items is checked after the fact | Catching drift and silent errors the checks miss |
| Two-person approval | One person prepares or accepts, a second approves | Anything your existing policies already require two signatures for |
Most workflows combine them. A typical progression is approval of everything at launch, then exception review plus sampling for the categories that measure well, with approval gates kept permanently for outbound and hard-to-reverse actions.
Route on signals you can trust
Exception review depends on deciding which items a person should see. It's tempting to ask the model how confident it is and route anything below a threshold. A model's own statement of confidence is not a reliable measure of whether it's right (research on this has found models tend to be overconfident when they state it), so use it, if at all, alongside signals your code can verify.
- Validation failures. A required field is missing, a part number isn't in the item master, a date is impossible, line totals don't add up.
- Business rules. Anything over a value threshold, a new supplier or customer, a first order for a part, a request that mentions a complaint or a legal term.
- Measured accuracy by category. Categories or fields that have scored poorly on your evaluation set always go to a person until they improve.
- Disagreement. Two extraction passes, or two models, that return different values for the same field.
- Unfamiliar input. A document layout, sender or language the workflow hasn't handled before.
Record why each item was routed to review and show that reason to the reviewer. "Part number not found in item master" focuses attention far better than a generic "needs review".
A reference design
Routing and approval are enforced in code; the model can't skip a queue by what it writes.
A review screen people will actually use
The review screen decides whether review is real. Treat it as a product in its own right, and test it with the people who'll use it.
Review screen essentials
- The source document or message side by side with the AI's output, with the relevant passage highlighted where possible.
- The reason the item was routed to review, stated plainly.
- Edit a single field in place, rather than rejecting and re-keying the whole record.
- A short list of rejection reasons, so rejections can be counted and grouped.
- Keyboard shortcuts for accept, edit and next, for high-volume queues.
- No "approve all" for items behind an approval gate, and no pre-ticked approvals.
- For actions, exactly what will happen on approval: the recipient, the record, the values before and after.
Watch for signs of rubber-stamping: approvals that take a second or two on items that need reading, or a reviewer who never edits anything. A simple countermeasure is to occasionally include a known-wrong item from your test set in the queue and check that it gets caught. Treat this as a check on the design, not on the people.
Approval gates for actions
When a workflow or an AI agent can do something, such as sending an email, creating an order or updating a customer record, the approval has to be enforced by your code, not requested in the prompt.
- Pause and queue. A gated action is stored with its exact payload and waits. Nothing runs until an approval is recorded.
- Approver authority from existing roles. Map who can approve what to the roles or groups you already manage in your identity provider, not to a separate list inside the AI tool.
- Specific and time-limited. An approval covers one action with one payload. If the payload changes, it needs approving again; if it sits too long, it expires.
- Re-validate at execution. Check the payload again when the approved action runs, in case the underlying record changed in the meantime.
- Separate requester and approver where money or policy is involved, as you would for the same action done by hand.
- Never approve on timeout. If nobody approves, the action waits or expires. It doesn't go ahead by default.
Our guide to putting AI agents to work on internal processes covers tool permissions and approval levels for agents in more detail, and securing internal apps with single sign-on and least privilege covers the role design approvals rely on.
Capture every decision and measure it
Every review is a labeled example. Capturing it is what lets the workflow improve and lets you reduce review safely.
What to log for each reviewed item
- The input, the AI's output, and the model and prompt version that produced it.
- Why the item was routed to review.
- The reviewer's decision, every edited field with before and after values, and the rejection reason if any.
- Who reviewed it, when, and how long it took.
Review measures worth tracking weekly
Accepted without edits
- What it tells you
- How often the AI's output is right as it stands, per category or field
Edit rate per field
- What it tells you
- Which fields need a better prompt, extraction step or validation rule
Review time per item
- What it tells you
- Whether review is actually faster than the manual process
Errors found in sampled items
- What it tells you
- Whether straight-through processing is as accurate as assumed
Queue age
- What it tells you
- Whether review is keeping up, or items are waiting too long
| Measure | What it tells you |
|---|---|
| Accepted without edits | How often the AI's output is right as it stands, per category or field |
| Edit rate per field | Which fields need a better prompt, extraction step or validation rule |
| Review time per item | Whether review is actually faster than the manual process |
| Errors found in sampled items | Whether straight-through processing is as accurate as assumed |
| Queue age | Whether review is keeping up, or items are waiting too long |
Add corrected items to your evaluation set, so the next prompt or model change is tested against the mistakes your reviewers actually found. Our guide to adding AI to document and email workflows covers that evaluation loop.
Reducing review without losing control
- Agree in advance what a category or field must achieve, over how many items, before it moves from full review to exception review. Write it down with the process owner.
- Move one category at a time, and keep sampling it at a meaningful rate afterward.
- Return a category to full review automatically if sampled errors or edit rates rise past an agreed level.
- Put everything back under full review for a period after any change to the model, prompt or extraction step, and graduate it again on the new results.
- Keep approval gates on outbound, financial and master-data actions permanently, whatever the accuracy.
Adding review to an AI workflow?
Tell us what the workflow produces and who checks it today. On a scoping call we'll talk through which outputs need review, where approval gates belong, what the review screen should show, and what a fixed-price build would cover.
Not ready to talk yet? Draft a project brief first
A one-page review design checklist
Agree these before launch
- Every output the workflow produces, what happens if it's wrong, and the review pattern for each.
- The signals that route an item to review, and the reason shown to the reviewer.
- Which actions sit behind an approval gate, and which roles can approve them.
- A review screen tested with the people who'll use it.
- What gets logged for each decision, and where it's stored.
- The measured results a category needs before its review is reduced, and the sampling rate after.
- The owner of the review queue and what happens when it backs up.
Our experience includes document and email processing, classification, summarization, automated workflows and AI agents, with integrations across the Anthropic Claude, OpenAI, Google Gemini and xAI Grok APIs. We design human review and approval steps into any workflow where output affects customers or money. See AI integration for how we'd approach your workflow, and compare your options if you're deciding whether to build it in-house or with outside help.
Sources
- Xiong et al., Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs (ICLR 2024)
- LLM06:2025 Excessive Agency (OWASP GenAI Security Project)
About the publisher
Published by Webb Technologies, drawing on 15+ years of professional software development. Our experience includes document and email processing, classification, summarization, automated workflows and AI agents with Claude, OpenAI, Grok and Gemini APIs. Round Table, our live multi-agent platform, runs at round-table.ai (opens in a new tab).
See it applied
- Demo with simulated data. Not client work.Try the AI workflow demo →Invoices read, checked and routed to a review queue, with thresholds you can move. It runs in your browser.
- Illustrative sample. Not client work.See a sample scope for an AI workflow →An invented supplier-invoice intake with a person reviewing low-confidence items and large amounts, on the company's own AI provider account and accepted on a labeled test set of its own invoices.
Share this guide
Related services
- AI integration AI built into your applications and business processes, with Claude, OpenAI, Gemini or Grok.
- Custom software development The review screens, approval queues and integrations around the model.
- AI Velocity Training Hands-on training so development teams adopt a shared, practical AI workflow.
Topics