On this page
In short
- Pick a process with real volume, a clear owner and examples where the right answer is already known.
- Measure today's baseline before building anything. Without it, "faster" and "more accurate" are opinions.
- Write the acceptance criteria, including the errors you won't tolerate, before the first prompt is written, and test against examples the build never saw.
- Run in shadow mode first, then with human review, and schedule the go or no-go decision on day one. Stopping is a legitimate outcome.
Why pilots stall
A pilot is an experiment, and an experiment needs a question it can answer. "Can AI help with supplier emails?" can't be answered; the honest reply is always "somewhat." "Can we classify incoming supplier emails into our eight queues with at least the accuracy of the current manual triage, while keeping a person on every exception?" can be.
The usual causes of a stalled pilot are predictable: a process chosen because it was exciting rather than measurable, no record of how the work performs today, success judged by how the demo felt, and no date on which someone decides. Each is fixed during scoping, not during the build.
Choosing the process
Signs of a good and a poor first AI pilot
Enough volume that small time savings add up
- Poor first pilot
- A task that happens a few times a month
Past examples where the correct answer is known
- Poor first pilot
- No way to say whether an output was right
A person already checks the work
- Poor first pilot
- Outputs would go straight to customers or suppliers unreviewed
Bounded inputs: one document type, one inbox, one form
- Poor first pilot
- Anything the business might send
Mistakes are caught and reversible
- Poor first pilot
- Errors affect safety, product release or control systems
A named process owner who wants it
- Poor first pilot
- A sponsor with no one doing the work day to day
| Good first pilot | Poor first pilot |
|---|---|
| Enough volume that small time savings add up | A task that happens a few times a month |
| Past examples where the correct answer is known | No way to say whether an output was right |
| A person already checks the work | Outputs would go straight to customers or suppliers unreviewed |
| Bounded inputs: one document type, one inbox, one form | Anything the business might send |
| Mistakes are caught and reversible | Errors affect safety, product release or control systems |
| A named process owner who wants it | A sponsor with no one doing the work day to day |
In a manufacturing company, candidates can include classifying and routing supplier or customer emails, extracting fields from certificates of analysis, packing slips or inspection reports, summarizing shift handover notes or maintenance logs, and answering questions from SOPs and work instructions with citations back to the source document. Each has volume, a known right answer and a person already in the loop.
Measure the baseline first
Before building anything, record how the process performs today. You'll need these numbers to judge the pilot, and collecting them often surfaces details that change the design.
Baseline measurements
- Volume: items per day or week, and how it varies.
- Handling time: how long a person spends per item, from a time study on a sample or from existing ticket or workflow timestamps.
- Turnaround: how long an item waits before it's handled.
- Error rate: how often today's process gets it wrong, from rework, returns, corrections or a sample re-checked by a second person.
- Error types: which mistakes are minor and which are costly.
- Current cost of the step, in hours rather than currency if that's easier to agree.
The error rate matters as much as the time. A pilot that's compared against a perfect human process will almost always look bad; one compared against the real process, including its mistakes, gets a fair hearing.
Write the acceptance criteria before you build
Acceptance criteria turn the pilot's question into numbers. Agree them with the process owner and whoever signs off on risk, write them into the pilot charter, and don't change them after seeing the results.
Types of acceptance criteria for an AI pilot
Quality
- How it's measured
- Accuracy on a held-out evaluation set, per category or field
- Example form
- At least the baseline error rate on every field that matters
Critical errors
- How it's measured
- Count of named, unacceptable mistakes
- Example form
- Zero cases where a wrong lot number passes review unflagged
Human effort
- How it's measured
- Review time per item with the AI's suggestion
- Example form
- Meaningfully less than the baseline handling time
Coverage
- How it's measured
- Share of items the system handles confidently
- Example form
- A target share, with the rest routed to people as today
Adoption
- How it's measured
- Share of suggestions accepted without edits
- Example form
- Tracked weekly during the pilot
Operational
- How it's measured
- Latency, cost per item, availability
- Example form
- Within limits the process can live with
| Criterion | How it's measured | Example form |
|---|---|---|
| Quality | Accuracy on a held-out evaluation set, per category or field | At least the baseline error rate on every field that matters |
| Critical errors | Count of named, unacceptable mistakes | Zero cases where a wrong lot number passes review unflagged |
| Human effort | Review time per item with the AI's suggestion | Meaningfully less than the baseline handling time |
| Coverage | Share of items the system handles confidently | A target share, with the rest routed to people as today |
| Adoption | Share of suggestions accepted without edits | Tracked weekly during the pilot |
| Operational | Latency, cost per item, availability | Within limits the process can live with |
Set the actual targets from your own baseline, not from a vendor's benchmark. Also define what counts as correct for each output and who judges it. Two reviewers disagreeing about the right answer is a finding in itself, and it's better discovered before the pilot than during it.
Build an evaluation set from real history
- Pull a representative sample of past items, including the awkward ones: scans, forwards, handwritten notes, unusual suppliers.
- Have the people who do the work record the correct answer for each.
- Split the set. Use one part while building and tuning; keep the other part aside, untouched, for the acceptance test.
- Score every version of the pilot against the same set, so changes to the prompt, model or extraction step are compared fairly.
- Add new hard cases found during the pilot to the building set, not the held-out one.
Our guide to adding AI to document and email workflows covers evaluation of a running pipeline in more detail, including tracking cost per document alongside accuracy.
Designing the pilot itself
A pilot should use real inputs and real people, but leave the existing process in charge until the criteria are met.
Nothing the pilot produces reaches a customer, supplier or production system without a person approving it.
- Shadow mode first. The AI runs on real items while people work exactly as before. Compare its output with theirs without affecting anything.
- Assisted mode second. Reviewers see the suggestion and accept, edit or reject it. Log every decision and how long it took.
- Tight scope. One process, one site or team, one kind of input, and a fixed pilot window agreed at the start.
- Security review up front. Confirm with IT security which data the pilot may send to a model provider, the provider's retention and training terms for your account type, and where logs are stored.
- Real integration, minimal footprint. Read the real inputs, but write results to the pilot's own review screen rather than back into the ERP or MES until the pilot passes.
If the process involves several steps that vary from case to case, rather than one classification or extraction, see putting AI agents to work on internal processes for the extra guardrails an agent needs.
The go or no-go decision
Put the decision meeting on the calendar when the pilot starts. Bring the criteria report, a sample of the worst errors and the reviewers' feedback. There are three good outcomes:
- Go. The criteria are met. Plan the production version: hardening, monitoring, a cost model at full volume, an owner, and the write-back into the system of record.
- Iterate once. Close, with a specific, named change that should fix it, such as a better extraction step for scanned documents. Rerun against the same held-out set.
- Stop. The criteria aren't met and no specific change looks likely to fix it. Record why. The evaluation set and baseline stay useful for the next attempt or the next process.
Planning a first AI pilot?
Tell us which process you're considering and how it runs today. On a scoping call we'll talk through whether it's a good first candidate, what the acceptance criteria could be, and what a fixed-price pilot would cover.
Not ready to talk yet? Draft a project brief first
A one-page pilot charter
Agree these before the build starts
- The process, the inputs in scope and the inputs explicitly out of scope.
- The process owner and the person who signs off on risk.
- Baseline volume, handling time, turnaround and error rate.
- Acceptance criteria, including the critical errors that fail the pilot outright.
- The evaluation set, who labeled it, and the held-out portion reserved for acceptance.
- Data handling approved by IT security, including the model provider's terms.
- The pilot window, the modes it runs in, and the date of the go or no-go decision.
Our experience includes document and email processing, classification, summarization, automated workflows and AI agents, with integrations across the Anthropic Claude, OpenAI, Google Gemini and xAI Grok APIs. Our own product, Round Table, runs Claude, ChatGPT, Grok and Gemini side by side in one conversation. See AI integration for how we'd approach a pilot, and how we work for the fixed-price process.
About the publisher
Published by Webb Technologies, drawing on 15+ years of professional software development. Our experience includes document and email processing, classification, summarization, automated workflows and AI agents with Claude, OpenAI, Grok and Gemini APIs. Round Table, our live multi-agent platform, runs at round-table.ai (opens in a new tab).
See it applied
- Demo with simulated data. Not client work.Try the AI workflow demo →Invoices read, checked and routed to a review queue, with thresholds you can move. It runs in your browser.
- Illustrative sample. Not client work.See a sample scope for an AI workflow →An invented supplier-invoice intake with a person reviewing low-confidence items and large amounts, on the company's own AI provider account and accepted on a labeled test set of its own invoices.
Share this guide
Related services
Topics