Reference
Glossary
Short, plain-language definitions of the terms used throughout this handbook.
| Term | Meaning |
|---|---|
| ANN | Approximate nearest neighbour — search that finds most of the closest vectors much faster than checking all of them |
| BM25 | A classic keyword-ranking formula used by search engines |
| Brute-force / flat search | Comparing the query with every stored vector; exact but slow at scale |
| Centroid | The centre point of a cluster, used by IVF |
| Chunk | A piece of a longer document, embedded separately |
| Collection | A group of records sharing vector size and metric (also called index, class or table) |
| Compaction | Background cleanup that removes deleted records and reorganises data |
| Cosine similarity | Similarity based on the angle between two vectors, from −1 to 1 |
| Cross-encoder / reranker | A model that scores a query and a document together for precise relevance |
| Dense vector | A vector where most values are non-zero, produced by embedding models |
| Dimensions | How many numbers a vector contains |
| DiskANN | A graph index designed to keep most data on SSD |
| Dot product (inner product) | Sum of the products of matching vector positions |
| efConstruction / efSearch | HNSW candidate-list sizes at build time and query time |
| Embedding | A vector produced by a model to represent meaning |
| Embedding model | A model that converts text, images or other data into embeddings |
| Euclidean distance (L2) | Straight-line distance between two points |
| Filter selectivity | The fraction of records that match a filter |
| HNSW | Hierarchical Navigable Small World — a layered graph index |
| Hybrid search | Combining vector search with keyword search |
| IVF | Inverted file index — clusters vectors and searches only the nearest clusters |
| k-means | A clustering algorithm that finds k cluster centres |
| kNN | k-nearest neighbours — the k closest vectors to a query |
| M | In HNSW, the maximum number of links per node |
| Metadata / payload | Extra fields stored with a vector (title, tenant, date…) |
| MRR | Mean reciprocal rank — how high the first relevant result appears |
| Multi-tenancy | Serving many customers from one system with data isolation |
| Namespace / partition | A sub-division of a collection, often per tenant |
| nDCG | A ranking metric that rewards putting the most relevant results first |
| nlist / nprobe | IVF's number of clusters, and number of clusters searched per query |
| Normalisation | Scaling a vector to length 1 |
| pgvector | A PostgreSQL extension that adds vector types and similarity search |
| Pre-filtering / post-filtering | Applying a filter before or after the vector search |
| Product quantization (PQ) | Compressing a vector by splitting it into pieces and storing codebook IDs |
| Quantization | Storing vectors with fewer bits to save memory |
| Recall@k | The fraction of the true top-k results that a search returned |
| Replica | A full copy of the index, used for throughput and availability |
| Rescoring | Re-ranking candidates found with compressed vectors using full-precision vectors |
| RRF | Reciprocal Rank Fusion — merges ranked lists using 1/(k + rank) |
| Scalar quantization (SQ) | Storing each number with fewer bits, e.g. 8-bit integers |
| Shard | A partition of the data stored on a separate node |
| Sparse vector | A mostly-zero vector over a vocabulary, used for keyword-style search |
| Tombstone | A marker that a record is deleted, before physical removal |
| Upsert | Insert a record, or update it if the ID already exists |
| Vector | An ordered list of numbers |
| Vector database | A database that stores vectors and finds the most similar ones quickly |