RAG Handbook retrieval-augmented generation Bipin Singh
The core pipeline

Embeddings

2 min readChapter 06 of 26By Bipin Singh

An embedding is a list of numbers — a vector — that represents the meaning of a piece of text. Texts with similar meaning get vectors that sit close together in space. This is the trick that makes semantic search possible: you can now measure how related two pieces of text are with arithmetic.

The intuition

Imagine every piece of text as a point in a high-dimensional space (hundreds or thousands of dimensions). The embedding model is trained so that "how do I reset my password" lands near "steps to recover account access," even though they share almost no words. Meaning, not keywords, determines closeness. That is why RAG can answer a question phrased completely differently from the source document.

Key idea
Keyword search matches words. Embedding search matches meaning. That's why a query and its answer can look nothing alike and still be found.

Dimensions and what they cost

An embedding's dimension is how many numbers are in the vector — commonly 384, 768, 1024, 1536, or 3072. Higher dimensions can capture more nuance but cost more to store and compare. Some modern models support Matryoshka embeddings, where you can truncate the vector to a shorter length and trade a little accuracy for a lot of storage and speed savings.

Choosing an embedding model

The choice matters as much as the LLM. Weigh:

TypeExamplesTrade-off
Hosted APIOpenAI text-embedding-3, Cohere Embed, Voyage, GeminiEasy, strong quality; per-token cost, data leaves your network
Open / self-hostedBGE, E5, GTE, Nomic, Jina, sentence-transformersPrivate, no per-call fee; you run the infrastructure

Model names and rankings change constantly — treat any specific recommendation as a snapshot and re-check the current benchmarks when you build.

Normalisation and distance

Many models return normalised vectors (length 1). When vectors are normalised, cosine similarity and dot product rank results identically, and dot product is cheaper. Check whether your model normalises; if not, normalise consistently for indexing and querying. See Retrieval for the distance metrics themselves.

Generating embeddings

// Index time: embed every chunk (batch for throughput).
const vectors = await embedder.embedBatch(chunks.map(c => c.text));
await db.upsert(chunks.map((c, i) => ({
  id: c.id,
  vector: vectors[i],
  text: c.text,
  metadata: c.metadata,
})));

// Query time: embed the question with the SAME model.
const queryVector = await embedder.embed(userQuestion);
Watch out
If you change embedding models, you must re-embed your entire corpus. Old and new vectors are not comparable — they live in different geometric spaces. Version your index so you never mix them.

Fine-tuning embeddings

When a general model can't tell your domain's near-synonyms apart, you can fine-tune an embedding model on pairs of (query, relevant passage) from your data. This pulls truly relevant pairs closer and pushes irrelevant ones apart, and can lift retrieval noticeably in specialised domains. It's an advanced step — exhaust chunking, hybrid search, and reranking first, since they're cheaper and often enough.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production RAG and LLM systems for enterprises.

Work with me