Reference
Cheat sheet
A one-page summary to come back to. Each line links to ideas explained earlier in the handbook.
Core ideas
- Embedding: a vector produced by a model so that similar meaning → nearby vectors.
- Vector database: stores vectors + metadata and finds nearest neighbours fast.
- Same model for indexing and querying. Store the model version with each vector.
Similarity
| Metric | Formula (vectors a, b) | Better when |
|---|---|---|
| Cosine similarity | a·b ÷ (‖a‖‖b‖) | higher |
| Dot product | Σ aᵢbᵢ | higher |
| Euclidean (L2) | √Σ(aᵢ − bᵢ)² | lower |
Normalised vectors → cosine = dot product, and all three give the same ranking.
Search
- Brute force: exact; cost ∝ vectors × dimensions. Fine up to roughly 100k vectors.
- ANN: approximate, much faster. Measure with recall@k vs exact search.
- Always report latency at a recall level.
Indexes
| Index | Idea | Key knobs | Notes |
|---|---|---|---|
| HNSW | Layered neighbour graph | M, efConstruction (build), efSearch (query) |
Default choice; memory-hungry |
| IVF | Clusters, search nearest few | nlist (build), nprobe (query) |
Needs training; low memory |
| PQ / SQ / BQ | Compress vectors | Bits / sub-vectors | 4–64× smaller; rescore to recover accuracy |
| DiskANN | Graph on SSD | — | Huge datasets, less RAM |
| Flat | Compare all | — | Exact; small or heavily filtered sets |
Tune order: query-time knob → filters/payload indexes → quantization → build params → hardware → model.
Memory
raw = vectors × dimensions × bytes per number
1M × 1,536 × 4 bytes ≈ 6.1 GB (int8: ≈ 1.5 GB)
Add index overhead, metadata and 30–50% headroom.
Using it well
- Stable IDs (
doc_id#chunk), provenance metadata, model version. - Filters: index filter fields; test filtered recall; enforce tenant/permission filters centrally.
- Hybrid: dense + keyword, fused with RRF (
Σ 1/(60 + rank)), then rerank the top ~50. - Updates: upsert new chunks, delete stale ones; deletes are tombstones until compaction.
- Model change: new index → backfill → evaluate → switch → delete old.
Choosing
- Already on Postgres and modest scale → pgvector.
- Already on Elasticsearch/OpenSearch, keyword-heavy → its vector search.
- Vector search is core, large scale, heavy filtering/multi-tenancy → dedicated vector database.
- Batch jobs or custom engines → FAISS / hnswlib.
- Prototype or local tool → embedded or brute force.
pgvector quick reference
CREATE EXTENSION vector;
ALTER TABLE docs ADD COLUMN embedding vector(1536);
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
SET hnsw.ef_search = 100;
SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10; -- <=> cosine, <-> L2, <#> negative inner product
Production checklist
Backups + tested restore · rebuild-from-source procedure · latency/memory/ingestion-lag monitoring · canary recall checks · auth, private networking, encryption · central permission filters · model-migration plan · load test with real filters.