RAG
Retrieval-augmented generation grounds a model's answers in your own data at query time, instead of relying only on what it learned during training.
- RAG (Retrieval-Augmented Generation) — RAG retrieves relevant documents at query time and feeds them to an LLM as context, so answers are grounded in your own data instead of just training data.
- Embeddings — an embedding turns text into a list of numbers positioned so similar meanings land near each other.
- Chunking — splits documents before you index them.
- Vector Database — finds the stored embeddings nearest a query.
- Reranking — a reranker re-scores the shortlist from your first search, reading query and document together.
- RAG Evaluation — scores retrieval quality and generation quality separately, because a RAG system can fail at either stage independently.