Skip to content
Webb Technologies

AI integration · Data access

Connecting AI to your internal data safely

Once a company starts using AI, an obvious next question is "can it answer questions from our own documents?": the procedures on SharePoint, the specs on a file share, the order history in the database. It can, and the usual technique, retrieval, is well understood. The hard part isn't getting answers; it's making sure each person only gets answers from data they were already allowed to see, from documents that are still current, with a source they can check. This guide covers how to do that.

Webb TechnologiesUpdated 10 min read

Drafted with AI assistance; facts checked against any sources cited. General guidance, not advice for your situation; verify before relying on it.

On this page

In short

  • Retrieval finds relevant passages and hands them to the model with the question. The model doesn't know who's asking, so permissions have to be enforced in the retrieval step, in code, before anything reaches it.
  • Carry the user's identity through: filter the index by the groups that can see each document, or query the source on the user's behalf so the source enforces its own permissions.
  • Index less than you can. Start with a curated collection that has an owner, exclude HR, legal and personal data by default, and fix overshared folders before they become searchable.
  • Keep the index in step with the sources, including deletions and permission changes, and prefer current approved versions over drafts and archives.
  • Give every answer citations people can open and check, and have the assistant say so when it can't find a source.

How retrieval works, and why permissions are the hard part

Retrieval, often called retrieval-augmented generation or RAG, works in two stages. Ahead of time, documents and records are split into passages and indexed, usually with both a keyword index and a vector index that matches on meaning. When someone asks a question, the system finds the most relevant passages and sends them to the model along with the question and instructions to answer only from them, with citations.

The model has no idea who is asking. Whatever passages it receives, it will use. So if the index was built by a service account that can read everything, and the search isn't filtered per user, the assistant becomes a friendly front end to the payroll folder. Telling the model "don't reveal HR data" in its instructions isn't a control; instructions can be ignored or talked around. The only reliable place to enforce access is before retrieval results reach the model.

Built-in assistants or a custom build

If your documents live in a platform that offers its own assistant, such as Microsoft 365 with Copilot, that may be the simplest route for general questions over documents. Microsoft states that Copilot only surfaces organizational data the signed-in user already has at least view permission for, which is exactly why cleaning up overshared sites comes first. Check current licensing, features and admin controls with the vendor.

When a custom retrieval build tends to make sense

Answers need data from databases or line-of-business systems

Why a built-in assistant may not cover it
Built-in assistants center on content in their own platform; connectors to other systems vary

Content is on on-premises file shares or several platforms

Why a built-in assistant may not cover it
Connectors vary, and permissions may not carry across

The answer feeds a specific workflow or screen

Why a built-in assistant may not cover it
A general chat window doesn't fit a quality or service process

You need tight control over sources, citations and logging

Why a built-in assistant may not cover it
You decide what's indexed, how it's ranked and what's recorded

You want to choose the model provider per workflow

Why a built-in assistant may not cover it
The platform decides which models its assistant offers and when they change

Carry permissions through to the user

There are a few established ways to make retrieval respect each user's access. They can be combined, and the right one depends on the source.

Patterns for permission-aware retrieval

Filter by access list

How it works
Each indexed passage stores the groups or users allowed to see its source. At query time, the user's group memberships from your identity provider are applied as a filter on the search
Watch for
Nested groups, sharing links and permission changes all need to be synced; stale access lists are a leak

Query on the user's behalf

How it works
The assistant searches the source system with the user's own delegated token, for example through Microsoft Graph, so the source enforces its permissions
Watch for
Permissions stay as current as the source's own search, but it's slower, depends on the source's search index, and is limited to what that search can do

Separate indexes per audience

How it works
One index per department, site or plant, and users only query the ones they belong to
Watch for
Simple and robust, but coarse; documents with mixed audiences don't fit neatly

Scoped database access

How it works
Queries go through narrow APIs or views that apply the user's scope, such as their plant or region, using row-level rules in the database or the API
Watch for
Never give the model a general-purpose SQL connection
  • Identity comes from sign-in, not from the prompt. The retrieval service takes the user from the validated SSO token, never from anything typed into the chat.
  • Deny by default. A passage with no access list, or one that failed to sync, is excluded rather than treated as public.
  • Check again at the source when it matters. For sensitive collections, confirm the user can still open the cited document before showing the answer.

Our guide to securing internal apps with single sign-on and least privilege covers group design and access tokens, and putting AI agents to work on internal processes covers narrow tools for when an assistant needs to look up records rather than read documents.

Decide what to index, and what not to

It's tempting to connect everything and let search sort it out. A smaller, curated start is safer and usually gives better answers, because the assistant isn't choosing between five versions of the same procedure.

Starting scope for an internal retrieval index

Approved policies and procedures

Exclude by default
HR, payroll and medical records

Work instructions and quality procedures

Exclude by default
Legal matters, contracts under negotiation, board material

Product specifications and datasheets

Exclude by default
Personal drives and mailboxes

Approved customer-service answers and FAQs

Exclude by default
Files containing credentials or connection strings

Engineering standards and templates

Exclude by default
Anything with a restricted sensitivity label
  • Give each collection an owner who decides what's in it and keeps it current.
  • Use the labels you have. If documents carry sensitivity labels or classification metadata, use them to exclude content automatically at ingestion.
  • Scan before indexing. Check for passwords, keys and personal data in files headed for the index, and hold back anything that matches for review.
  • Index structured data carefully. For database records, it can be better to look up live values through a narrow API at question time than to copy rows into an index where they go stale.

Stale and conflicting data

An index is a copy, and copies drift. A superseded procedure that stays in the index will be quoted with the same confidence as the current one, and a document someone deleted or restricted yesterday may still be searchable today.

  • Sync changes, deletions and permissions. Use the source's change feed where it has one, such as delta queries in Microsoft Graph, and agree a maximum lag. Permission changes and deletions should propagate at least as fast as new content.
  • Prefer the approved version. Index published versions, not drafts; drop or clearly tag archives; and record each document's effective date and status as metadata.
  • Rank by currency where it matters. When two passages conflict, filter or rank on status and date in code rather than relying on the model to notice.
  • Show the date. Every citation includes the document's date or revision, so readers can spot an old source.
  • Review what gets cited. Reports of the most-cited documents show owners which content the business actually relies on, and which needs updating.

Citations people can check

Citations turn an answer from something to trust into something to verify. They also make errors visible quickly, because a reader can see when a cited passage doesn't say what the answer claims.

What each answer should show

  • Links to the source documents or records, opening in the source system with the user's own permissions.
  • The passage or section the answer relied on, not just the file name.
  • Each source's date or revision.
  • A plain "I couldn't find this in the approved sources" when retrieval finds nothing relevant, rather than an answer from the model's general knowledge.
  • A way to flag a wrong or outdated answer, routed to the collection's owner.

Least privilege for the pipeline itself

The ingestion and search services hold a copy of a lot of company data, so they need the same care as any other system that does.

  • Read-only, narrowly scoped connectors. The ingestion identity can read only the sites, folders or views being indexed. On Microsoft 365, permissions such as Graph's Sites.Selected allow access to specific sites rather than all of them.
  • Database logins limited to views. A read-only login with access to specific views, not the tables behind them.
  • Treat the index as sensitive as its sources. Passages and their vector representations are derived from your documents; encrypt the store, restrict access to it, and include it in backups and retention decisions.
  • Protect the logs. Questions and retrieved passages can contain sensitive data. Log them for audit and quality, with access limited to named roles.
  • Treat indexed content as untrusted input. A document can contain text written to manipulate the model. The assistant should never gain permissions or take actions because of something it read.
  • Confirm provider terms. Check what the model provider may retain from prompts that now include your documents, under your account's current terms.

A reference design

internal-retrieval / permission-awarediagram
Read-only connectors pull approved content from SharePoint, file shares and database views. Ingestion filters out excluded and sensitive content, splits documents into passages and stores each passage with its access list, date and status. At question time the query service takes the user's identity from SSO, filters the search to what that user may see, sends the matching passages to the model, and returns an answer with citations. Every question, result and citation is logged.

Permissions are applied before the model sees anything, so no instruction or question can widen access.

Test access, not just answers

Answer quality is tested like any AI workflow, with real questions and expected sources. Access needs its own tests, run before launch and after every change to connectors, groups or the index.

  • Test as different people. Use test accounts in different groups, plants and roles, and confirm each sees only the sources it should.
  • Plant marker documents. Put a restricted test document containing a unique phrase in each sensitive area, then ask about that phrase as users who shouldn't see it. It should never appear in an answer or a citation.
  • Test removals. Remove a test user from a group, or delete a marker document, and measure how long until the assistant stops using it.
  • Test hostile content. Index a test document containing instructions aimed at the model and confirm they have no effect.

Our guide to scoping a first AI pilot covers building an evaluation set from real history and writing acceptance criteria before you build.

Want AI to answer from your own documents and data?

Tell us where the content lives, who should see what, and what questions people need answered. On a scoping call we'll talk through what to index first, how permissions would carry through, and what a fixed-price build would cover.

Request a scoping call

Not ready to talk yet? Draft a project brief first 

A one-page checklist

Before connecting AI to internal data

  • Broad and outdated permissions reviewed and fixed on every source being connected.
  • A curated first collection with a named owner; HR, legal, personal and restricted content excluded by default.
  • Permissions enforced in retrieval code using the signed-in user's identity, denying by default.
  • Connectors and database logins read-only and scoped to what's indexed.
  • Changes, deletions and permission updates synced within an agreed lag.
  • Drafts and archives excluded or tagged; dates and status stored with every passage.
  • Citations with links, passages and dates on every answer, and a clear "not found" response.
  • Index, logs and backups protected like the source data; provider terms confirmed.
  • Access tested with multiple test users and marker documents, before launch and after changes.

Our experience includes document and email processing, classification, summarization, automated workflows and AI agents, with integrations across the Anthropic Claude, OpenAI, Google Gemini and xAI Grok APIs, along with OAuth and OIDC authentication flows and database design across SQL Server, PostgreSQL and NoSQL stores. Before rolling an assistant out to staff, agree an AI acceptable-use policy that sets the rules for what data goes where. See AI integration for how we'd approach connecting AI to your data.

Sources

Want to talk through your version of this?

A 30-minute call. You leave with a clear approach and the real risks, whether or not you hire us.