Quantization — making vectors smaller
Vectors are big. A single 1,536-dimension embedding stored as 32-bit floats takes 6 KB; a hundred million of them need roughly 600 GB just for the raw numbers. Quantization compresses vectors so more fit in memory and distance calculations get cheaper, at a small cost in accuracy.
Where the bytes go
bytes per vector = dimensions × bytes per number
1,536 dims × 4 bytes (float32) = 6,144 bytes ≈ 6 KB
1 million vectors ≈ 6.1 GB
100 million vectors ≈ 614 GB
Memory is often the biggest cost of running a vector database, because the fastest indexes want vectors in RAM.
Scalar quantization (SQ)
Store each number with fewer bits. The most common form maps every float to an 8-bit integer (int8), using the range of values seen in the data:
float32: 0.8134, -0.2210, 0.0457 ... (4 bytes each)
int8: 104, -28, 6 ... (1 byte each)
- Saving: 4× smaller.
- Accuracy: usually a very small drop for typical text embeddings.
- Effort: simple, no complex training.
Some systems also support 16-bit floats (half precision) — 2× smaller with almost no accuracy loss.
Product quantization (PQ)
PQ compresses much harder by splitting each vector into chunks and replacing each chunk with the ID of its closest "prototype":
- Split every vector into
msub-vectors (e.g. 1,536 dims → 96 chunks of 16 dims). - For each chunk position, learn 256 prototype sub-vectors with k-means — a codebook.
- Store each chunk as the ID (0–255, one byte) of its nearest prototype.
original: 1,536 floats = 6,144 bytes
PQ (m = 96): 96 one-byte codes = 96 bytes → 64× smaller
Distances are computed with precomputed lookup tables, which is also fast. The trade-off is a larger accuracy loss than scalar quantization, and PQ needs training data.
Binary quantization (BQ)
The most extreme option: keep only the sign of each number — 1 bit per dimension.
[0.81, -0.22, 0.04, -0.67] → [1, 0, 1, 0]
1,536 dims → 1,536 bits = 192 bytes → 32× smaller
Comparisons become bit operations (Hamming distance), which are extremely fast. Accuracy depends heavily on the embedding model — some models are designed to work well with binary quantization; others lose a lot.
Rescoring: get the accuracy back
The standard trick is a two-step search:
- Search the compressed vectors to find, say, the top 100 candidates quickly.
- Rescore those 100 using the original full-precision vectors (kept on disk or in cheaper storage) and return the true top 10.
This recovers most of the lost accuracy while keeping memory low — the compressed index sits in RAM, the full vectors don't need to.
Comparing the options
| Method | Size vs float32 | Accuracy impact | Training needed | Good default when… |
|---|---|---|---|---|
| float16 | ½ | Negligible | No | You want an easy, safe saving |
| Scalar (int8) | ¼ | Small | Minimal | Most production workloads |
| Product (PQ) | 1/16 – 1/64 (tunable) | Moderate | Yes | Very large datasets, memory-bound |
| Binary | 1/32 | Model-dependent | No | Huge scale with a compatible model, plus rescoring |
Quantization trades a little accuracy for a lot of memory. With rescoring, you can often keep most of the accuracy and still cut memory by 4× or more.
Another way to shrink vectors is to use fewer dimensions. Some embedding models are trained so that their vectors can be truncated (often called Matryoshka embeddings) — check whether yours supports it before reaching for heavier compression.