Most companies have the same problem: the answer exists somewhere in a PDF, a SharePoint folder or a shared drive, and nobody can find it. An AI assistant that searches your documents and answers in plain language, with a link to where it found the answer, is one of the most useful and least risky first AI projects you can do.
The technique is called retrieval-augmented generation, or RAG. The name is complicated; the idea isn’t.
How it works
- Collect: point the system at your documents: a storage bucket, SharePoint, OneDrive, Google Drive, Confluence or your intranet.
- Index: it splits each document into passages and turns each passage into a list of numbers (an “embedding”) that captures its meaning. These go into a search index.
- Ask: when someone asks a question, the system finds the passages closest in meaning.
- Answer: a language model writes an answer using only those passages, and shows which documents it used.
Because the model answers from your passages instead of its general knowledge, answers are more accurate and can be checked. When the answer isn’t in your documents, a well-built system says so.
First, check whether you need to build anything
If your documents live in Microsoft 365 or Google Workspace, the assistant that comes with your suite may already do this. Microsoft 365 Copilot and Gemini in Google Workspace can both answer questions from files the user has access to. See our post on Copilot or Gemini. Build your own when you need answers inside your own app or website, for customers, for documents outside those suites, or with more control over sources and behaviour.
The building blocks on each cloud
| AWS | Azure | Google Cloud | |
|---|---|---|---|
| Managed RAG | Amazon Bedrock Knowledge Bases | Azure AI Search with Microsoft Foundry | Vertex AI Search, or Vertex AI RAG Engine |
| Models | Many providers’ models through Amazon Bedrock | OpenAI and other models through Microsoft Foundry | Gemini and other models through Vertex AI |
| Typical sources | S3, SharePoint, OneDrive, Google Drive, Confluence, web crawler | Blob Storage, SharePoint, databases | Cloud Storage, Google Drive, websites, databases |
Microsoft renamed Azure AI Foundry to Microsoft Foundry, so you will see both names in guides. On AWS, Bedrock Knowledge Bases handles the parsing, embeddings, search index and citations for you, and now offers agentic retrieval for questions that need reasoning across several documents.
What it costs to run
- Indexing: a one-off cost to process your documents, plus small top-ups when they change.
- The search index: often the biggest fixed cost. Some search services bill every hour whether or not anyone asks a question. Pick the size for your document volume, not the largest tier.
- Questions: each answer uses the language model, billed by the amount of text read and written, or per query on some managed services.
- Storage for the documents themselves, usually small.
For an internal assistant over a few thousand documents, the running cost is usually modest compared with the time staff spend searching. Ask for an estimate based on your document count and expected questions per day before you start.
Where your data stays
- Your documents and index stay in the region you choose. All three clouds have regions in India.
- The model may not be available in every region. Check which models run in your chosen Indian region. Some services can route requests to other regions for capacity (AWS calls this cross-region inference), which may process data outside India; switch it off if that matters to you.
- Training: the three providers state that business customers’ prompts and data in these services aren’t used to train their foundation models. Read the terms for the specific model you choose.
- Permissions: if different staff may see different documents, the assistant must respect that. Plan this from the start; it is the most common gap in quick prototypes.
Questions to answer before you build
- Who will use it: staff, customers, or both?
- Which documents, and who keeps them up to date?
- Must answers respect who can see which document?
- Must the data stay in India?
- What should it say when it doesn’t know?
- How will you measure whether it helps? (For example: questions answered without a human, time to answer.)
A good first project
Start with one set of documents and one group of users: HR policies for staff, product manuals for the support team, or standard contracts for sales. Build it in a few weeks, collect 50 real questions, check the answers with the people who own the documents, then widen it. Small and checked beats large and untrusted.
An example: HR policies for staff
A 300-person company keeps about 80 HR documents: leave rules, travel policy, reimbursements, benefits. HR answers the same questions every week. A document assistant indexes the current versions, answers staff questions inside Teams or the intranet with a link to the policy, and says “please contact HR” when the answer isn’t in the documents. HR reviews a sample of answers each month and updates the policies the assistant struggled with.
Keeping it up to date
An assistant is only as good as the documents behind it. Give each document set an owner, sync the index automatically when files change, remove old versions instead of keeping them “just in case”, and review a handful of real questions and answers every month.
Mistakes we see
- Feeding it everything. Old drafts and outdated policies produce confident wrong answers. Curate first.
- No citations. Without sources, people can’t check answers and stop trusting them.
- Ignoring permissions. One shared index for everyone can leak documents meant for a few.
- The largest search tier “to be safe”, paid every hour for a few hundred documents.
Where DevOps TechLab fits
We build document assistants on AWS, Azure and Google Cloud, in an Indian region, with citations and permissions designed in from the start. Most projects start with a two-to-four-week pilot on one set of documents, so you see real answers to real questions before deciding to go further.
Questions people ask
What is RAG?
Retrieval-augmented generation: an AI system finds the passages in your documents most relevant to a question, then a language model writes an answer from those passages and shows its sources.
Which cloud is best for a document assistant?
All three have managed options: Amazon Bedrock Knowledge Bases, Azure AI Search with Microsoft Foundry, and Vertex AI Search or RAG Engine. The best fit is usually the cloud where your documents and team already are.
Can the data stay in India?
Your documents and index can stay in an Indian region on all three clouds. Check that your chosen model is available there, and turn off cross-region routing if data must not leave India.
Do we need to build one if we have Copilot or Gemini?
Not always. If your documents are in Microsoft 365 or Google Workspace and staff are the users, Copilot or Gemini may be enough. Build your own for customers, other document stores or more control.
What does it cost?
Indexing once, a search index (often the biggest fixed cost), and a charge per question for the language model. For a few thousand internal documents the running cost is usually modest.
Have documents nobody can find answers in?
Tell us what they are and who asks the questions. We will suggest a two-to-four-week pilot and estimate the running cost.







