Sizing and scaling
Before you deploy, estimate how much memory and compute you'll need. A few lines of arithmetic prevent nasty surprises in your cloud bill or in production latency.
Estimating memory
Start with the raw vectors:
raw bytes = number of vectors × dimensions × bytes per number
| Vectors | Dimensions | float32 (4 bytes) | int8 (1 byte) |
|---|---|---|---|
| 1 million | 768 | ~3.1 GB | ~0.8 GB |
| 1 million | 1,536 | ~6.1 GB | ~1.5 GB |
| 10 million | 768 | ~30.7 GB | ~7.7 GB |
| 10 million | 1,536 | ~61.4 GB | ~15.4 GB |
Then add:
- Index overhead. HNSW stores neighbour links for every vector — with
M = 16, roughly 100–200 bytes extra per vector on top of the vector itself (more with higherM). IVF adds very little. - Metadata. Text snippets and payload fields can easily outweigh small vectors. Measure an average record.
- Headroom. Leave room for growth, updates, compaction and the operating system — often 30–50%.
Estimate with a real sample: load 100,000 records, check the memory or disk your database reports, and scale up linearly.
Reducing the footprint
| Lever | Typical effect |
|---|---|
| Scalar quantization (int8) | ~4× less vector memory |
| Fewer dimensions (if your model supports truncation) | Proportional saving |
| Store full vectors on disk, quantized in RAM, rescore | Large RAM savings with small accuracy loss |
| Keep large text out of the vector database | Smaller records; fetch details from your main store |
| Disk-based indexes (e.g. DiskANN) | Most data on SSD instead of RAM |
Scaling throughput: replicas
If one node can hold the data but can't serve enough queries per second, add replicas — full copies of the index — and spread queries across them. Replicas also give you high availability: if one node fails, others keep serving.
Scaling data size: shards
If the data doesn't fit on one node, split it into shards, each holding part of the vectors. A query is sent to every shard, each returns its local top k, and the results are merged.
query → shard 1 → top 10 ┐
→ shard 2 → top 10 ├→ merge → global top 10
→ shard 3 → top 10 ┘
Shards add latency (you wait for the slowest one) and coordination overhead, so shard only when you need to. Sharding by tenant can avoid fan-out entirely when queries always target one tenant.
Cost drivers
- Memory — usually the biggest factor for in-RAM indexes.
- Embedding generation — re-embedding large corpora isn't free; cache and hash to avoid repeats.
- Query compute — high
efSearch/nprobeand rerankers cost CPU per query. - Replicas — each one multiplies memory cost.
- Managed service pricing — compare on your actual volumes; pricing models differ (per pod/node, per storage, per query).
Size from arithmetic, confirm with a real sample, and scale in this order: tune and compress first, add replicas for throughput, add shards for size.