Skip to content
Webb Technologies

AI integration · Pilots

Scoping a first AI pilot with acceptance criteria you can measure

Most first AI projects don't fail loudly. They demo well, run for a while, and then sit in a state nobody can call finished or stopped, because nobody agreed what success would look like. This guide covers how to scope a first pilot so it ends in a decision: which process to pick, what to measure before you start, how to write acceptance criteria that can be checked, and how to run the pilot without putting real work at risk.

Webb TechnologiesUpdated 7 min read

Drafted with AI assistance; facts checked against any sources cited. General guidance, not advice for your situation; verify before relying on it.

On this page

In short

  • Pick a process with real volume, a clear owner and examples where the right answer is already known.
  • Measure today's baseline before building anything. Without it, "faster" and "more accurate" are opinions.
  • Write the acceptance criteria, including the errors you won't tolerate, before the first prompt is written, and test against examples the build never saw.
  • Run in shadow mode first, then with human review, and schedule the go or no-go decision on day one. Stopping is a legitimate outcome.

Why pilots stall

A pilot is an experiment, and an experiment needs a question it can answer. "Can AI help with supplier emails?" can't be answered; the honest reply is always "somewhat." "Can we classify incoming supplier emails into our eight queues with at least the accuracy of the current manual triage, while keeping a person on every exception?" can be.

The usual causes of a stalled pilot are predictable: a process chosen because it was exciting rather than measurable, no record of how the work performs today, success judged by how the demo felt, and no date on which someone decides. Each is fixed during scoping, not during the build.

Choosing the process

Signs of a good and a poor first AI pilot

Enough volume that small time savings add up

Poor first pilot
A task that happens a few times a month

Past examples where the correct answer is known

Poor first pilot
No way to say whether an output was right

A person already checks the work

Poor first pilot
Outputs would go straight to customers or suppliers unreviewed

Bounded inputs: one document type, one inbox, one form

Poor first pilot
Anything the business might send

Mistakes are caught and reversible

Poor first pilot
Errors affect safety, product release or control systems

A named process owner who wants it

Poor first pilot
A sponsor with no one doing the work day to day

In a manufacturing company, candidates can include classifying and routing supplier or customer emails, extracting fields from certificates of analysis, packing slips or inspection reports, summarizing shift handover notes or maintenance logs, and answering questions from SOPs and work instructions with citations back to the source document. Each has volume, a known right answer and a person already in the loop.

Measure the baseline first

Before building anything, record how the process performs today. You'll need these numbers to judge the pilot, and collecting them often surfaces details that change the design.

Baseline measurements

  • Volume: items per day or week, and how it varies.
  • Handling time: how long a person spends per item, from a time study on a sample or from existing ticket or workflow timestamps.
  • Turnaround: how long an item waits before it's handled.
  • Error rate: how often today's process gets it wrong, from rework, returns, corrections or a sample re-checked by a second person.
  • Error types: which mistakes are minor and which are costly.
  • Current cost of the step, in hours rather than currency if that's easier to agree.

The error rate matters as much as the time. A pilot that's compared against a perfect human process will almost always look bad; one compared against the real process, including its mistakes, gets a fair hearing.

Write the acceptance criteria before you build

Acceptance criteria turn the pilot's question into numbers. Agree them with the process owner and whoever signs off on risk, write them into the pilot charter, and don't change them after seeing the results.

Types of acceptance criteria for an AI pilot

Quality

How it's measured
Accuracy on a held-out evaluation set, per category or field
Example form
At least the baseline error rate on every field that matters

Critical errors

How it's measured
Count of named, unacceptable mistakes
Example form
Zero cases where a wrong lot number passes review unflagged

Human effort

How it's measured
Review time per item with the AI's suggestion
Example form
Meaningfully less than the baseline handling time

Coverage

How it's measured
Share of items the system handles confidently
Example form
A target share, with the rest routed to people as today

Adoption

How it's measured
Share of suggestions accepted without edits
Example form
Tracked weekly during the pilot

Operational

How it's measured
Latency, cost per item, availability
Example form
Within limits the process can live with

Set the actual targets from your own baseline, not from a vendor's benchmark. Also define what counts as correct for each output and who judges it. Two reviewers disagreeing about the right answer is a finding in itself, and it's better discovered before the pilot than during it.

Build an evaluation set from real history

  1. Pull a representative sample of past items, including the awkward ones: scans, forwards, handwritten notes, unusual suppliers.
  2. Have the people who do the work record the correct answer for each.
  3. Split the set. Use one part while building and tuning; keep the other part aside, untouched, for the acceptance test.
  4. Score every version of the pilot against the same set, so changes to the prompt, model or extraction step are compared fairly.
  5. Add new hard cases found during the pilot to the building set, not the held-out one.

Our guide to adding AI to document and email workflows covers evaluation of a running pipeline in more detail, including tracking cost per document alongside accuracy.

Designing the pilot itself

A pilot should use real inputs and real people, but leave the existing process in charge until the criteria are met.

ai-pilot / shadow and assisted modesdiagram
Real items from the existing inbox, folder or system are copied to the pilot. The AI step produces suggestions that go to a review screen; reviewers accept, edit or reject them. The existing process stays the system of record, and every suggestion, decision and timing is logged for measurement against the acceptance criteria.

Nothing the pilot produces reaches a customer, supplier or production system without a person approving it.

  • Shadow mode first. The AI runs on real items while people work exactly as before. Compare its output with theirs without affecting anything.
  • Assisted mode second. Reviewers see the suggestion and accept, edit or reject it. Log every decision and how long it took.
  • Tight scope. One process, one site or team, one kind of input, and a fixed pilot window agreed at the start.
  • Security review up front. Confirm with IT security which data the pilot may send to a model provider, the provider's retention and training terms for your account type, and where logs are stored.
  • Real integration, minimal footprint. Read the real inputs, but write results to the pilot's own review screen rather than back into the ERP or MES until the pilot passes.

If the process involves several steps that vary from case to case, rather than one classification or extraction, see putting AI agents to work on internal processes for the extra guardrails an agent needs.

The go or no-go decision

Put the decision meeting on the calendar when the pilot starts. Bring the criteria report, a sample of the worst errors and the reviewers' feedback. There are three good outcomes:

  • Go. The criteria are met. Plan the production version: hardening, monitoring, a cost model at full volume, an owner, and the write-back into the system of record.
  • Iterate once. Close, with a specific, named change that should fix it, such as a better extraction step for scanned documents. Rerun against the same held-out set.
  • Stop. The criteria aren't met and no specific change looks likely to fix it. Record why. The evaluation set and baseline stay useful for the next attempt or the next process.

Planning a first AI pilot?

Tell us which process you're considering and how it runs today. On a scoping call we'll talk through whether it's a good first candidate, what the acceptance criteria could be, and what a fixed-price pilot would cover.

Request a scoping call

Not ready to talk yet? Draft a project brief first 

A one-page pilot charter

Agree these before the build starts

  • The process, the inputs in scope and the inputs explicitly out of scope.
  • The process owner and the person who signs off on risk.
  • Baseline volume, handling time, turnaround and error rate.
  • Acceptance criteria, including the critical errors that fail the pilot outright.
  • The evaluation set, who labeled it, and the held-out portion reserved for acceptance.
  • Data handling approved by IT security, including the model provider's terms.
  • The pilot window, the modes it runs in, and the date of the go or no-go decision.

Our experience includes document and email processing, classification, summarization, automated workflows and AI agents, with integrations across the Anthropic Claude, OpenAI, Google Gemini and xAI Grok APIs. Our own product, Round Table, runs Claude, ChatGPT, Grok and Gemini side by side in one conversation. See AI integration for how we'd approach a pilot, and how we work for the fixed-price process.

Want to talk through your version of this?

A 30-minute call. You leave with a clear approach and the real risks, whether or not you hire us.