Skip to content
Webb Technologies

Illustrative sample, not a client project

This scope is for an invented project, written to show what the document you’d receive contains.

Illustrative sample · written scope

Sample scope of work: an AI invoice workflow.

The document you review before accepting a fixed price. This one is for an invented AI invoice intake: fields read into an existing accounting system, a person reviewing the exceptions, and acceptance measured on a labeled test set.

Email to a colleague

Jump to the document ↓

Illustrative sample, not a client project

The client, company, systems, suppliers and names in this document are invented. It shows the structure and level of detail of the real document; it is not a record of work for a client.

Written scope · fixed-price project

Supplier invoice intake with AI and human review

Client
Example Parts Distribution Co. (fictional)
Prepared by
Webb Technologies
Document
Scope of work and fixed-price basis
Status
Illustrative sample, not a client project

1.Background and goals

Example Parts Distribution Co. (“the Client”) is a mid-sized distributor of industrial parts. Supplier invoices arrive as PDF attachments, some of them scans, in a shared accounts payable mailbox. An accounts payable (AP) clerk opens each one, checks it against the purchase order and keys the header and line items into the Client's accounting system by hand. Invoices wait in the mailbox at busy times, keying errors are found late, and the same invoice is occasionally entered twice.

This project delivers an AI workflow that reads each invoice, extracts its fields, checks them against the Client's own data and writes the result into the accounting system's existing invoice import, with a person reviewing anything the workflow is unsure of and anything above an amount threshold. It runs on the Client's own AI provider account. The goals are:

  • AP clerks review and correct extracted invoices on one screen instead of keying them from the PDF.
  • Low-confidence items, anything that fails a check and any invoice above the amount threshold go to a person before they reach the accounting system.
  • Accuracy is measured before go-live on a labeled test set of the Client's own past invoices, against thresholds agreed in phase 1, and measured again whenever a prompt or model changes.
  • Every extraction, check, review decision and correction is recorded, so AP and Finance can see what the AI did and why.
  • The Client's IT team can run, change and extend the system after handover without us.

2.In scope

  • Mailbox intake. New messages in the shared AP mailbox are picked up automatically, and each PDF attachment becomes one item in a processing queue. Messages without a usable attachment go to the review queue with the reason.
  • Extraction. An AI model on the Client's own provider account reads each invoice, including scanned pages, and returns a fixed set of fields: supplier, invoice number, invoice date, due date, purchase order number, currency, subtotal, tax, freight, total, and for each line the description, supplier part number, quantity, unit price and line total. A field that isn't on the invoice is returned empty, never guessed.
  • Checks. Each item is checked against the Client's own data: the supplier matches the supplier list, the purchase order exists and is open for that supplier, the lines add up to the total, the total matches the purchase order within the tolerance agreed in phase 1, and the supplier and invoice number haven't been received before.
  • Review routing. An item goes to a person when any field is below its confidence threshold, any check fails, the total is above the amount threshold, the supplier is new or unmatched, the invoice mentions new bank or remittance details, or the document isn't an invoice. Thresholds are agreed in phase 1 from the test results (section 7); the amount threshold and tolerances are administrator settings, not code.
  • Review screen. A web screen where AP clerks see the invoice next to the extracted fields, with the fields that need attention highlighted, and approve, correct or reject each item with a reason. A second approval by the AP lead for items above the amount threshold.
  • Write-back. Approved items, and items that pass every check without review, are written to the accounting system's invoice staging table, where the Client's existing import picks them up as unapproved invoices. Approval for payment stays in the accounting system, as today.
  • Test set and scoring. A test set of the Client's own past invoices, redacted and labeled with the AP team, and a scoring run that reports accuracy per field and per supplier group, the review rate and the silent error rate (items that were wrong and would have gone through without review).
  • Regression checks. The full test set runs in the CI/CD pipeline whenever the prompt, the extraction schema or the model version changes, and blocks the change if any field falls below its agreed threshold.
  • Audit trail. Every item's extracted values, check results, review decisions, corrections, and the prompt and model versions that produced it, recorded with the user and time.
  • Sign-in. Single sign-on through the Client's existing identity provider, over OpenID Connect or SAML, with three roles taken from groups the Client's IT team manages: AP reviewer, AP lead and administrator.
  • Usage tracking. AI usage recorded per invoice, with an alarm when daily usage passes a limit the Client sets.
  • Infrastructure and delivery. Infrastructure defined as code (AWS CDK), deployed through a CI/CD pipeline (GitHub Actions) into separate test and production environments.
  • Monitoring. Alarms for processing failures, AI provider errors and timeouts, a growing review queue, write-back failures and sign-in failures.
  • Documentation and handover. Runbooks, a data dictionary, architecture and security documentation, and knowledge-transfer sessions, as set out in section 11.

3.Out of scope

  • Approving invoices for payment, paying suppliers or matching goods received. The workflow prepares invoices for the accounting system; approval and payment stay with the Client's existing processes.
  • Any change to supplier records, including bank and remittance details. Invoices that mention new details are flagged for a person and never update anything.
  • Credit notes, statements, quotes and other document types. They are recognized and sent to the review queue unprocessed; extracting them is a candidate for a later phase, quoted separately.
  • Any mailbox other than the shared AP mailbox, and invoices that arrive on paper or through a supplier portal.
  • Training or fine-tuning AI models. The workflow uses the chosen provider's models as they are, with prompts, a fixed output schema and the Client's own data for checks.
  • Any write to the accounting system other than the invoice staging table.
  • Mobile apps and offline use.
  • Integration with any other business system.
  • Hardware, network changes, software licenses and AI provider subscriptions. We specify what is needed; the Client's teams make the changes and hold the licenses and accounts.
  • Support, maintenance or new features after handover. Available as separately quoted work.

4.Systems and data sources

Systems this project connects to, the access it needs and who owns each one on the Client's side

System
Client's email system: shared AP mailbox
Role in this project:
Source of incoming invoices
Access needed:
Read access to this one mailbox only
Owner (Client):
IT
System
AI provider account (Client-owned: Anthropic Claude, OpenAI, Google Gemini or xAI Grok, chosen in phase 1)
Role in this project:
Reads each invoice and returns the extracted fields
Access needed:
An API key for this workflow only, with usage limits where the provider supports them; retention and training settings reviewed in phase 1
Owner (Client):
IT and Finance
System
SQL Server: accounting database
Role in this project:
Supplier list, open purchase orders and invoices already received, for the checks
Access needed:
Read-only login to three named views, reached from the Client's AWS account over its site-to-site VPN
Owner (Client):
IT (database administrator)
System
SQL Server: invoice staging table
Role in this project:
Where approved invoices are written for the accounting system's existing import
Access needed:
A login that can insert into the staging table only
Owner (Client):
IT (database administrator)
System
Identity provider (Client's existing)
Role in this project:
Sign-in and roles for the review screen
Access needed:
An SSO app registration and three groups
Owner (Client):
IT
System
Past invoices (test set source)
Role in this project:
Source for the labeled test set
Access needed:
A sample of past invoices with their keyed values, exported by AP and redacted as agreed with IT security
Owner (Client):
Finance (AP lead)
System
AWS account (Client-owned)
Role in this project:
Hosts the workflow, review screen, stored invoices and monitoring
Access needed:
A deployment role for the pipeline; named user access for us during the project
Owner (Client):
IT
System
GitHub organization (Client-owned)
Role in this project:
Repositories, CI/CD pipeline and the test set
Access needed:
Member access for us during the project
Owner (Client):
IT
Architecture overview (illustrative): invoices arrive in a shared mailbox on the Client's email system, staff sign in through the Client's identity provider, the workflow runs in the Client's AWS account and calls the Client's own AI provider account, and results reach the accounting system only through the staging table, reached over the site-to-site VPN. The full design, with every connection and network rule, is prepared in phase 1.

5.Assumptions and dependencies

  • The Client provides the access in section 4 before the build phase starts. If access arrives later, the schedule moves with it; the price changes only if the scope does.
  • The Client has, or opens, a business account with one of Anthropic Claude, OpenAI, Google Gemini or xAI Grok. Which provider and model to use is decided in phase 1 by running the tuning part of the test set against the candidates the Client is willing to use. AI usage is billed to that account.
  • The Client's IT security team confirms in phase 1 which data may be sent to the chosen provider, and which retention and training settings apply under the Client's own agreement with it.
  • The AP team provides the past invoices for the test set, and AP clerks label the expected answers with us using a written labeling guide. Two clerks label a shared subset independently so ambiguous rules are found and settled.
  • Finance and the AP lead sign off, in phase 1, the per-field accuracy thresholds, the acceptable review rate and silent error rate, the amount threshold and the purchase order tolerance, based on the baseline and the first test results.
  • The accounting system's existing import from the staging table works and is documented. Changes to that import are the Client's.
  • The supplier list and purchase orders in the accounting database are correct. Where they aren't, the Client corrects them in the accounting system; the workflow doesn't keep its own copy.
  • A site-to-site VPN (or AWS Direct Connect) links the Client's network to its AWS account, set up by the Client's network team.
  • The Client names one product owner who can answer questions and accept deliverables, and one contact each in AP and IT security.
  • The test environment uses the redacted test set and a copy of the accounting views with test data. No unredacted production invoices are used in test.
  • Hosting, license and other third-party costs are billed to the Client's own accounts.

6.Deliverables

  • Source code for the workflow and the review screen, with automated tests, in the Client's repositories.
  • The prompts and the output schema, versioned in the Client's repositories alongside the code.
  • The redacted test set, the answer key, the labeling guide and the scoring script, in the Client's repositories under the same access controls as the source data.
  • Database migration scripts for the workflow's own tables, run by the pipeline, so the schema is versioned with the code.
  • Infrastructure as code (AWS CDK) for the test and production environments.
  • A CI/CD pipeline (GitHub Actions) that tests, builds and deploys every change, runs the test set when a prompt, the schema or the model version changes, and authenticates to AWS with OIDC so no AWS keys are stored.
  • An architecture and data-flow document for the Client's security review, including every network rule with its source, destination, port and purpose, and exactly what is sent to the AI provider.
  • A data dictionary covering every extracted field with its format and normalization rule, and every table, view and field the workflow reads or writes.
  • Test reports: the phase 1 baseline, each scored run during the build, and the final run on the held-out part of the test set.
  • Runbooks: deploy, roll back, rotate credentials (including the AI provider key), change a threshold or the amount limit, run the test set and read its report, add corrected items to the test set, upgrade the model version, and respond to each alert.
  • Knowledge-transfer sessions, listed in section 11.

7.Acceptance criteria

The system is accepted when every criterion below passes in production. Each one is checked together with the Client's product owner.

Accuracy criteria are measured on the held-out part of the test set: invoices from the Client's own history, labeled with the AP team, that are never used while prompts are being tuned. Each field is scored separately after normalization, so a better average can't hide a worse total or date.

Acceptance criteria and how each one is checked

Ref
AC1
Criterion:
Accuracy for each extracted field meets the threshold agreed for that field in phase 1.
How it is checked:
A scored run of the held-out test set with the production prompt, schema and model version, reported per field and per supplier group and reviewed with the AP lead.
Ref
AC2
Criterion:
With the review rules applied, the silent error rate (items that were wrong and would have gone through without review) is at or below the level agreed in phase 1.
How it is checked:
The same run, scored with the production routing rules and thresholds.
Ref
AC3
Criterion:
The share of items routed to review is within the level the AP lead agreed the team can handle in phase 1.
How it is checked:
The same run; the review rate reported alongside AC2.
Ref
AC4
Criterion:
Every invoice with a total above the amount threshold, from a new or unmatched supplier, or mentioning new bank or remittance details goes to a person, and no field the answer key marks as absent is filled in without review.
How it is checked:
Every such item in the held-out set checked in the scored run.
Ref
AC5
Criterion:
Malformed model output, a provider timeout or outage, a corrupt or encrypted attachment, the same invoice received twice, instructions hidden in an invoice and a failed write to the staging table each end with the item in the review queue or waiting to retry, never partly written or written twice.
How it is checked:
One test case per failure, run in test.
Ref
AC6
Criterion:
A change to the prompt, the schema or the model version that lowers any field below its threshold is blocked by the pipeline.
How it is checked:
A deliberately worse prompt change submitted in test and shown to be blocked with its report.
Ref
AC7
Criterion:
Every item's extracted values, check results, review decisions, corrections and prompt and model versions appear in its audit trail with the user and time.
How it is checked:
The audit trail compared with the actions taken during the AC1 run and a set of test reviews.
Ref
AC8
Criterion:
Nothing in the workflow can write to the accounting database except by inserting into the staging table.
How it is checked:
IT reviews both logins' permissions; any other write attempted with either login is refused.
Ref
AC9
Criterion:
Usage is recorded per invoice, and the usage alarm fires when the daily limit is passed.
How it is checked:
A lowered limit in test and a batch of test invoices.
Ref
AC10
Criterion:
The Client's team deploys a change and rolls it back using only the runbooks.
How it is checked:
Done during knowledge transfer, observed by both parties.

Accuracy, review-rate and silent-error thresholds, the amount threshold and the purchase order tolerance: stated here in the real document, agreed with the Client in phase 1 from the baseline and the first test results.

Acceptance is judged against these criteria only. Anything not listed here is not a reason to withhold acceptance, and is not in the price; it can be raised as a change (section 9).

8.Milestones

Work runs in five phases. Each phase ends on a written exit criterion, and scored test runs and working software are reviewed with the Client in the test environment throughout phases 2 and 3.

Project phases and the exit criterion that closes each one

Phase
1. Access, test set and design
What happens:
Access set up. Past invoices sampled across suppliers, layouts and scans, including non-invoices, duplicates and high-value items; redacted, labeled with AP and split into tuning and held-out parts. The current manual process measured as a baseline. Provider and model chosen on the tuning set. Fields, checks, review rules and the data design drafted and reviewed with AP, Finance and IT.
Exit criterion:
Written sign-off of the answer key by the AP lead, of the thresholds and amount threshold by Finance, and of the design and the data sent to the provider by the Client's IT lead.
Phase
2. Extraction and checks
What happens:
Intake, extraction, checks and review routing built, and scored against the tuning set after each change.
Exit criterion:
A scored run on the tuning set reviewed with the AP lead; AC5 passes in test.
Phase
3. Review screen and write-back
What happens:
Review screen, second approval, write-back to the staging table, usage tracking and the pipeline's regression run built and reviewed in test with AP clerks. Sign-in and roles in place.
Exit criterion:
Product owner approves the review screen; AC6, AC8 and AC9 pass in test.
Phase
4. Acceptance and go-live
What happens:
Final scored run on the held-out set, then deployment to production. For an agreed period, every invoice is reviewed by a person whatever the routing says, and the results are compared before the review rules take effect.
Exit criterion:
All acceptance criteria pass in production.
Phase
5. Handover
What happens:
Documentation delivered, knowledge-transfer sessions held, access handed over.
Exit criterion:
Handover checklist signed.

Milestone dates: stated here in the real document, agreed with the Client before work starts.

9.Change process

  1. Either party raises a change in writing: what is needed and why.
  2. We reply in writing with its effect on scope, price and schedule, or confirm that it has none.
  3. Nothing changes until the Client's product owner approves it in writing. Approved changes are added to this document as numbered amendments.
  4. A change that isn't approved isn't built, and the original scope and price stand.

Until the system is accepted, fixing something that fails an acceptance criterion in section 7 is not a change. It is covered by the fixed price.

10.Security and IT review

The following are prepared for the Client's IT and security teams during phase 1, before the workflow connects to any Client system or sends any invoice to the AI provider:

  • A data-flow diagram showing every connection and its direction, including exactly what is sent to the AI provider: the invoice pages and the instructions for the fixed set of fields, and nothing from the accounting database.
  • A network rule list: source, destination, port, protocol and purpose for each rule. The only connection into the Client's network is the workflow's SQL Server connection over the site-to-site VPN, on a single network rule the Client's network team controls.
  • The AI provider's retention and training settings for the Client's account, reviewed with IT security under the Client's own agreement with the provider.
  • Database logins: one read-only login for the accounting views and one that can only insert into the staging table, named and documented. No shared or personal credentials in the running system.
  • The model only returns fields in a fixed schema. It has no access to the Client's systems and cannot send email, write data or call other services; every write is made by the workflow's own code after the checks.
  • Text in an invoice, such as instructions to ignore rules or pay a different account, is treated as content to extract, never as an instruction. Invoices mentioning new bank or remittance details always go to a person.
  • Sign-in only through the Client's existing identity provider, so the Client's existing multi-factor and access policies apply. Roles come from groups the Client's IT team manages in that identity provider.
  • Secrets, including the AI provider API key, kept in AWS Secrets Manager, never in code, configuration files, tickets or email.
  • Encryption in transit on every connection, and at rest for stored invoices, the test set, extracted data and backups.
  • Logging of sign-ins, deployments, provider errors and application errors to CloudWatch, retained according to the Client's policy.
  • Dependency scanning in the pipeline, with findings reported to the Client before each release.
  • Data classification: supplier names and addresses, invoice details, purchase order references, and the names and work email addresses of AP staff. The test set is redacted with consistent placeholders and kept under the same access controls as the source invoices.
  • Answers to the Client's security questionnaire, and an architecture walkthrough for the security team.

11.Handover

Handover is part of the fixed price. It is complete when the Client's team has deployed a change, rolled it back and responded to a test alert themselves, using only the documentation. It includes:

  • Everything in section 6, confirmed in the Client's repositories and accounts.
  • Knowledge-transfer sessions: architecture walkthrough, code tour, supervised deployment, rollback and recovery, running the test set and reading its report, changing a threshold, and a test incident.
  • Access and credentials: every account confirmed as the Client's, secrets rotated, and our access removed or reduced to whatever support is agreed.

The full contents of a handover package are shown in the sample handover package. It was written for a different invented project, a plant-floor dashboard, but it is organized the same way for a workflow like this one.

12.Ownership

  • The Client owns all code, infrastructure definitions, documentation and data produced under this scope. There is no license fee and no lock-in.
  • Everything is built in the Client's repositories and AWS account. Nothing runs in accounts we control.
  • Open-source components remain under their own licenses, which are listed in each repository.
  • Confidentiality is covered by the NDA signed before this engagement.

13.Price and what it covers

Fixed price: stated here in the real document.

The fixed price covers everything in sections 2, 6, 7 and 11: design, build, testing, deployment to test and production, documentation, knowledge transfer, handover, and fixing anything that fails an acceptance criterion before acceptance.

It does not cover:

  • Anything listed as out of scope in section 3.
  • Approved changes under section 9, each priced before it starts.
  • Hosting, software licenses and other third-party costs, billed directly to the Client's accounts.
  • Support, maintenance or new features after handover, quoted separately if wanted.

Invoicing schedule and payment terms: stated here in the real document.

14.Approval

Nothing is built until this scope is accepted. Signing below accepts the scope and the fixed price stated in section 13.

For the Client

Name
Signature
Date

For Webb Technologies

Name
Signature
Date

Use it as a checklist

What to look for in a scope for an AI workflow.

Whoever you hire, a scope worth signing answers these questions in writing before work starts.

  1. Accuracy is measured on your own documents.

    Acceptance is judged on a labeled test set of your past documents, not on demo files, with part of it held back for the final run.

  2. The thresholds are agreed before the build.

    Accuracy per field, how much goes to review and the amount above which a person always checks are set with the people who own the process, from test results.

  3. It says what goes to a person, and why.

    Low-confidence fields, failed checks, large amounts and anything unusual reach a person before they reach your system.

  4. It covers what happens when the AI fails.

    Bad output, provider outages, duplicates and instructions hidden in a document each have a tested path.

  5. Prompt and model changes are retested.

    A change to the prompt, the schema or the model version reruns the test set and is blocked if an important field gets worse.

  6. The price section says what isn't covered.

    Hosting, licenses and other third-party costs, changes and support after launch are listed, so nothing outside the price comes as a surprise. AI usage is billed to your own provider account.

Keep reading

The companion sample.

The sample handover is for a different invented project, a plant-floor dashboard. It's organized the same way for any system.

Want a scope like this for your AI workflow?

A 30-minute call. If it's a fit, one to two weeks of working sessions, then a scope like this one and one fixed price before any build work starts.