Vector Database Handbook from zero to production Bipin Singh
Choosing & running

Sizing and scaling

2 min readChapter 19 of 25By Bipin Singh

Before you deploy, estimate how much memory and compute you'll need. A few lines of arithmetic prevent nasty surprises in your cloud bill or in production latency.

Estimating memory

Start with the raw vectors:

raw bytes = number of vectors × dimensions × bytes per number
Vectors Dimensions float32 (4 bytes) int8 (1 byte)
1 million 768 ~3.1 GB ~0.8 GB
1 million 1,536 ~6.1 GB ~1.5 GB
10 million 768 ~30.7 GB ~7.7 GB
10 million 1,536 ~61.4 GB ~15.4 GB

Then add:

Tip

Estimate with a real sample: load 100,000 records, check the memory or disk your database reports, and scale up linearly.

Reducing the footprint

Lever Typical effect
Scalar quantization (int8) ~4× less vector memory
Fewer dimensions (if your model supports truncation) Proportional saving
Store full vectors on disk, quantized in RAM, rescore Large RAM savings with small accuracy loss
Keep large text out of the vector database Smaller records; fetch details from your main store
Disk-based indexes (e.g. DiskANN) Most data on SSD instead of RAM

Scaling throughput: replicas

If one node can hold the data but can't serve enough queries per second, add replicas — full copies of the index — and spread queries across them. Replicas also give you high availability: if one node fails, others keep serving.

Scaling data size: shards

If the data doesn't fit on one node, split it into shards, each holding part of the vectors. A query is sent to every shard, each returns its local top k, and the results are merged.

query → shard 1 → top 10 ┐
      → shard 2 → top 10 ├→ merge → global top 10
      → shard 3 → top 10 ┘

Shards add latency (you wait for the slowest one) and coordination overhead, so shard only when you need to. Sharding by tenant can avoid fan-out entirely when queries always target one tenant.

Cost drivers

Key idea

Size from arithmetic, confirm with a real sample, and scale in this order: tune and compress first, add replicas for throughput, add shards for size.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me