RAG Handbook retrieval-augmented generation Bipin Singh
Building & reference

Glossary

2 min readChapter 25 of 26By Bipin Singh

Quick definitions for the vocabulary used throughout these docs.

RAG
Retrieval-Augmented Generation — retrieving relevant text at query time and giving it to an LLM to ground its answer.
Embedding
A vector of numbers representing the meaning of text, so similarity can be measured mathematically.
Vector
An ordered list of numbers; here, the output of an embedding model.
Dimension
The number of values in a vector (e.g. 768, 1536).
Chunk
A passage a document is split into for indexing and retrieval.
Chunk overlap
Text repeated between adjacent chunks so meaning isn't lost at boundaries.
Vector database
A store optimised for fast nearest-neighbour search over vectors, with metadata filtering.
ANN
Approximate Nearest Neighbour — fast search that trades a little accuracy for large speed gains.
HNSW
A graph-based ANN index prized for its speed/recall balance.
Cosine similarity
A similarity measure based on the angle between two vectors.
Top-k
The number of chunks retrieved for a query.
Dense retrieval
Semantic search using embeddings (matches meaning).
Sparse retrieval
Keyword search such as BM25 (matches exact terms).
Hybrid search
Combining dense and sparse retrieval and fusing the results.
RRF
Reciprocal Rank Fusion — merging ranked lists by position rather than raw score.
MMR
Maximal Marginal Relevance — balancing relevance against diversity in results.
Reranking
A second, more accurate pass that reorders retrieved candidates by relevance.
Cross-encoder
A model that scores a (query, document) pair together; used for reranking.
Bi-encoder
A model that embeds query and document separately; used for retrieval.
HyDE
Hypothetical Document Embeddings — embedding a generated draft answer to improve retrieval.
Faithfulness
Whether an answer's claims are supported by the retrieved context.
Hallucination
A confident, fluent, unsupported (often false) model output.
Grounding
Tying an answer to provided source context.
Prompt injection
Malicious instructions that hijack the model; "indirect" when hidden in retrieved content.
GraphRAG
RAG over a knowledge graph of entities and relationships.
Agentic RAG
A system that treats retrieval as a tool and reasons in a loop across sources.
Contextual retrieval
Prepending situating context to each chunk before embedding it.
Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me