retrieval-augmented generation
Letting a model look up relevant documents first, then answer using them.
A language model only knows what it saw in training, which may be out of date or missing your private information. With retrieval-augmented generation, or RAG, the system first searches your documents for passages related to the question. It adds those passages to the prompt, and the model answers from them. This keeps answers current and reduces hallucination.
Think of it as
Like an open-book exam. Instead of answering from memory, the student first finds the right pages, then writes the answer from what is in front of them.
Example
An employee asks the company assistant how many leave days they can carry over. The system finds the current HR policy page, passes it to the model, and the answer quotes this year's rule instead of a guess.