On this page
In short
- Start with narrow, checkable tasks: sorting, pulling fields out, summarizing. Leave decisions and outbound actions to people at first.
- Treat every email and attachment as untrusted input. The model's output is a suggestion your system validates, not a command it follows.
- Choose a provider by testing it on your own documents, and check its API data-use terms and hosting options before any real data flows.
- Build an evaluation set from past documents before launch, and rerun it whenever you change the prompt or the model.
Where AI fits in a document workflow
Large language models are good at reading messy, inconsistent text and producing something structured. That maps onto four jobs that show up in almost every back office:
Four common AI tasks in document and email workflows
Classification
- Example
- Is this email an order, a complaint, an invoice query or spam?
- How it's checked
- Against a fixed list of categories
Extraction
- Example
- Pull PO number, part numbers, quantities and dates from a PDF
- How it's checked
- Against a schema and your master data
Summarization
- Example
- Condense a long supplier thread for the buyer who picks it up
- How it's checked
- By the person reading it
Drafting
- Example
- Propose a reply or a routing note for a person to edit and send; see designing human review into AI workflows
- How it's checked
- By the person sending it
| Task | Example | How it's checked |
|---|---|---|
| Classification | Is this email an order, a complaint, an invoice query or spam? | Against a fixed list of categories |
| Extraction | Pull PO number, part numbers, quantities and dates from a PDF | Against a schema and your master data |
| Summarization | Condense a long supplier thread for the buyer who picks it up | By the person reading it |
| Drafting | Propose a reply or a routing note for a person to edit and send; see designing human review into AI workflows | By the person sending it |
What these have in common is that the output can be checked, by code, by a person, or both. Tasks where the model makes a decision nobody reviews, or takes an action outside your systems, belong later, once you've seen how it performs on your documents.
A reference design
The model is one step in a pipeline, not the pipeline. Most of the engineering, and most of the safety, is in the steps around it.
The model never writes to a system of record directly; validated output does.
- Intake. Read from a shared mailbox through Microsoft Graph or your mail platform's API, or from a monitored folder or bucket. Use an app registration or service account scoped to that mailbox or folder only; Exchange Online, for example, can limit an app's mail permissions to named mailboxes.
- Extraction. Get clean text before the model sees it. Many PDFs contain real text; scanned ones need OCR. Recent models can also read images and PDFs directly, which is worth testing on your worst documents.
- Structured output. Ask for a specific JSON shape. The major providers each offer a way to request output that matches a schema, which removes a whole class of parsing errors.
- Validation. Check the output with ordinary code: required fields present, part numbers that exist in your item master, dates that make sense, totals that add up.
- Routing. Results that pass go straight through; anything uncertain or failing goes to a person.
Human review that actually gets used
"Human in the loop" only works if the review step is faster than doing the work by hand. Design the review screen with the same care as the automation.
- Show the source document next to the extracted fields, with the relevant passage highlighted where possible.
- Let reviewers correct a field in place rather than re-keying the record.
- Save every correction. Corrections are the best test cases you'll ever get.
- Start in suggestion mode, where people approve everything, and let categories graduate to straight-through processing only once their measured accuracy justifies it.
Data handling and security
Your security team will ask where the data goes, who can see it, and what the provider does with it. Have the answers written down before the first real document is sent.
Data handling checklist
- Use the provider's paid or business API tier, not a consumer chat app or unpaid tier, and read its current terms on training use and data retention for API traffic.
- Decide whether to call the provider directly or through your cloud: Claude is available through Amazon Bedrock and Google Cloud Vertex AI, Gemini through Vertex AI, and OpenAI models through Microsoft Azure. That can keep traffic, billing and access control inside an account you already govern.
- Send only what the task needs. Strip signatures, disclaimers and unrelated attachments; redact fields the model doesn't need to see.
- Store API keys in a secrets manager, rotate them, and keep separate keys per environment.
- Log inputs and outputs for audit and debugging, with the same retention and access rules as the source documents.
- Confirm with legal or compliance whether any document types, such as HR, medical or export-controlled material, must stay out of scope.
Choosing between Claude, OpenAI, Gemini and Grok
We've integrated all four of these providers' APIs. Our honest view is that the right model depends on your documents, and that the differences between providers change often enough that any ranking written today will be stale soon. What stays true:
- Test on your documents. Run the same evaluation set through two or three candidates. Scanned forms, long email threads and spreadsheets-as-PDFs expose differences quickly.
- Weigh hosting and terms alongside quality. Availability through a cloud you already use, regional options and contractual terms can matter more than a small accuracy difference.
- Use the smallest model that passes. Many classification and extraction tasks run well on faster, cheaper model tiers; save the largest models for hard summarization or reasoning.
- Keep a thin abstraction. Put the provider call behind one internal interface so you can switch models or run two side by side without rewriting the workflow.
Our own product, Round Table, puts Claude, ChatGPT, Grok and Gemini in one conversation, with streaming responses and usage tracking (token accounting) built in. Building it is where much of our view on provider differences and usage comes from.
Evaluate before launch and after every change
A demo on five hand-picked emails tells you nothing about the five hundred you'll see next month. Evaluation is what turns "it seems to work" into a number you can manage.
- Collect a representative sample of past documents, including the ugly ones: forwards, replies, scans and multi-language messages.
- Record the correct answer for each, ideally from the people who do the work today.
- Measure accuracy per category and per field, not just overall. A single field that's often wrong can sink a workflow that looks fine on average.
- Rerun the set whenever you change the prompt, the model or the extraction step, and compare against the last run before releasing.
- Add every correction from the review queue to the set, so it grows with the real mix of documents.
- Track cost per document alongside accuracy, so a model change that's slightly better but much more expensive is a visible decision.
Have a document queue in mind?
Tell us which inbox or document type is eating your team's time. On a scoping call we'll talk through the pipeline, the review step and what your security team will ask, and what a fixed-price build would cover.
Not ready to talk yet? Draft a project brief first
A sensible rollout
Pick one document type with clear volume and a clear owner. Run the pipeline in shadow mode alongside the current process and compare. Move to suggestion mode with review of everything, then let the categories that measure well go straight through. Only then consider the next document type, reusing the same intake, validation, review and logging pieces.
Our experience includes document and email processing, classification, summarization and automated workflows with AI. If the steps change from case to case, putting AI agents to work on internal processes covers when an agent fits better than a pipeline. For how we'd approach yours, see AI integration; if your developers want to use AI tools in their own work, AI Velocity Training is the related offer.
Sources
- Role Based Access Control for Applications in Exchange Online (Microsoft)
- PDF and file input: Claude, OpenAI, Gemini
- Structured output: Claude, OpenAI, Gemini, xAI
- Cloud hosting: Claude on Amazon Bedrock, Claude on Vertex AI, models on Vertex AI, Azure OpenAI
- LLM01: Prompt injection (OWASP GenAI Security Project)
About the publisher
Published by Webb Technologies, drawing on 15+ years of professional software development. Our experience includes document and email processing, classification, summarization, automated workflows and AI agents with Claude, OpenAI, Grok and Gemini APIs. Round Table, our live multi-agent platform, runs at round-table.ai (opens in a new tab).
See it applied
- Demo with simulated data. Not client work.Try the AI workflow demo →Invoices read, checked and routed to a review queue, with thresholds you can move. It runs in your browser.
- Illustrative sample. Not client work.See a sample scope for an AI workflow →An invented supplier-invoice intake with a person reviewing low-confidence items and large amounts, on the company's own AI provider account and accepted on a labeled test set of its own invoices.
Share this guide
Related services
Topics