Submit a Story

Glossary

Retrieval-Augmented Generation

A technique that finds relevant passages in your own data and adds them to the prompt, so a language model answers from those sources instead of from memory alone.

Retrieval-augmented generation, usually shortened to RAG, combines search with a language model. When a question comes in, the system first retrieves the most relevant pieces of content from a knowledge base, then puts them in the prompt and asks the model to answer using them.

A typical pipeline splits documents into pieces (document chunking), turns them into embeddings, stores them in a vector database and finds matches with semantic search. The retrieved text lands in the context window next to the question.

RAG is the standard way to give a model current or private information without retraining it, and it reduces LLM hallucination by AI agent grounding. Its quality depends mostly on retrieval: if the wrong passages come back, the answer will be wrong too, so test the search step on its own.

Related: Large Language Model, AI Agent.

← All terms