Large language models are remarkably capable, but on their own they don't know your products, policies or internal documents. Ask them about your business and they'll either say they don't know or, worse, make something up.
Retrieval-augmented generation, or RAG, solves this by giving the model the right information at the moment it answers.
How RAG works
- Your documents are split into small passages and indexed so they can be searched by meaning.
- When someone asks a question, the system retrieves the most relevant passages.
- Those passages are given to the language model along with the question.
- The model writes an answer grounded in that content, ideally citing its sources.
Where RAG works well
- Internal knowledge assistants for policies, procedures and documentation.
- Customer support answers based on help-centre articles and product guides.
- Searching large collections of contracts, reports or technical documents.
- Onboarding assistants that help new staff find information quickly.
Where it needs care
RAG is only as good as the content it retrieves. Outdated, contradictory or poorly structured documents lead to poor answers. Permissions also matter: users should only receive answers drawn from content they're allowed to see.
Before building, ask: is our knowledge base accurate, current and organised? Cleaning it up is often the most valuable step.
RAG versus fine-tuning
Fine-tuning changes a model's behaviour or style by training it on examples. RAG supplies up-to-date facts at question time. For most knowledge assistants RAG is simpler, cheaper to update and easier to audit, because you can see which sources were used.
Measuring quality
A responsible RAG project includes evaluation: a set of real questions with known good answers, tested whenever prompts, models or content change. Track answer accuracy, citation quality, response time and cost, and give users an easy way to escalate to a person.