Vector Database Handbook from zero to production Bipin Singh
How it works inside

Quantization — making vectors smaller

2 min readChapter 08 of 25By Bipin Singh

Vectors are big. A single 1,536-dimension embedding stored as 32-bit floats takes 6 KB; a hundred million of them need roughly 600 GB just for the raw numbers. Quantization compresses vectors so more fit in memory and distance calculations get cheaper, at a small cost in accuracy.

Where the bytes go

bytes per vector = dimensions × bytes per number
1,536 dims × 4 bytes (float32) = 6,144 bytes ≈ 6 KB

1 million vectors   ≈ 6.1 GB
100 million vectors ≈ 614 GB

Memory is often the biggest cost of running a vector database, because the fastest indexes want vectors in RAM.

Scalar quantization (SQ)

Store each number with fewer bits. The most common form maps every float to an 8-bit integer (int8), using the range of values seen in the data:

float32: 0.8134, -0.2210, 0.0457 ...   (4 bytes each)
int8:      104,     -28,      6 ...     (1 byte each)

Some systems also support 16-bit floats (half precision) — 2× smaller with almost no accuracy loss.

Product quantization (PQ)

PQ compresses much harder by splitting each vector into chunks and replacing each chunk with the ID of its closest "prototype":

  1. Split every vector into m sub-vectors (e.g. 1,536 dims → 96 chunks of 16 dims).
  2. For each chunk position, learn 256 prototype sub-vectors with k-means — a codebook.
  3. Store each chunk as the ID (0–255, one byte) of its nearest prototype.
original: 1,536 floats = 6,144 bytes
PQ (m = 96): 96 one-byte codes = 96 bytes   → 64× smaller

Distances are computed with precomputed lookup tables, which is also fast. The trade-off is a larger accuracy loss than scalar quantization, and PQ needs training data.

Binary quantization (BQ)

The most extreme option: keep only the sign of each number — 1 bit per dimension.

[0.81, -0.22, 0.04, -0.67]  →  [1, 0, 1, 0]
1,536 dims → 1,536 bits = 192 bytes   → 32× smaller

Comparisons become bit operations (Hamming distance), which are extremely fast. Accuracy depends heavily on the embedding model — some models are designed to work well with binary quantization; others lose a lot.

Rescoring: get the accuracy back

The standard trick is a two-step search:

  1. Search the compressed vectors to find, say, the top 100 candidates quickly.
  2. Rescore those 100 using the original full-precision vectors (kept on disk or in cheaper storage) and return the true top 10.

This recovers most of the lost accuracy while keeping memory low — the compressed index sits in RAM, the full vectors don't need to.

Comparing the options

Method Size vs float32 Accuracy impact Training needed Good default when…
float16 ½ Negligible No You want an easy, safe saving
Scalar (int8) ¼ Small Minimal Most production workloads
Product (PQ) 1/16 – 1/64 (tunable) Moderate Yes Very large datasets, memory-bound
Binary 1/32 Model-dependent No Huge scale with a compatible model, plus rescoring
Key idea

Quantization trades a little accuracy for a lot of memory. With rescoring, you can often keep most of the accuracy and still cut memory by 4× or more.

Tip

Another way to shrink vectors is to use fewer dimensions. Some embedding models are trained so that their vectors can be truncated (often called Matryoshka embeddings) — check whether yours supports it before reaching for heavier compression.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me