The data model
Different vector databases use different names, but almost all share the same basic model. Once you know it, switching between products is mostly a matter of vocabulary.
The building blocks
| Concept | Also called | What it is |
|---|---|---|
| Collection | Index, class, table | A group of records that share a vector size and similarity metric |
| Record | Point, object, row, document | One stored item |
| ID | Key, primary key | A unique identifier you choose (or the database generates) |
| Vector | Embedding | The list of numbers used for similarity search |
| Metadata | Payload, properties, attributes, columns | Extra fields: text, tags, dates, numbers, IDs |
| Namespace | Partition, tenant, shard key | A sub-division inside a collection, often per customer |
A single record might look like this:
{
"id": "kb-1042#chunk-3",
"vector": [0.012, -0.044, 0.318, "... 1,533 more numbers ..."],
"metadata": {
"doc_id": "kb-1042",
"title": "Why was my payment refused?",
"section": "Card payments",
"language": "en",
"updated_at": "2026-08-14",
"tenant_id": "acme",
"source_url": "https://example.com/help/payments/refused",
"text": "If your card is declined at a shop, the most common reasons are..."
}
}
Designing records well
Choose stable IDs. Derive IDs from your source data — document ID + chunk number — so re-running ingestion updates records instead of creating duplicates.
Store what you filter on. Any field you'll want to filter by (tenant, language, date, category, access group) must be in the metadata — and usually indexed. See Metadata filtering.
Store what you display, or a pointer to it. Search results need something to show. Either keep the text snippet in metadata, or keep an ID and fetch details from your main database. Keeping text in the vector database is simpler; keeping it elsewhere avoids duplication and size limits.
Record provenance. Keep the source document ID, the chunk position, and the embedding model name/version. You'll need them for updates, deletes, citations and re-embedding.
The vector finds the record; the metadata makes the result useful — for filtering, display, permissions and maintenance.
One collection or several?
| Use separate collections when… | Use one collection with metadata when… |
|---|---|
| Vectors come from different embedding models or have different sizes | Same model and size |
| Data has completely different access patterns or lifecycles | Items are searched together |
| You need hard isolation (e.g. regulated customers) | Logical separation by filter is enough |
Multiple vectors per item
Sometimes one item deserves several vectors — a product with a text description and a photo, or a document with a title vector and a body vector. Options:
- Separate records that share a
doc_id, combined after search. - Named vectors — some databases let one record hold several vectors, each searchable on its own.
Where the source of truth lives
Treat the vector database as a search index, not your system of record. Keep original documents and business data in your primary database or storage, and design ingestion so the vector index can always be rebuilt from them. This makes model changes, migrations and recovery far less scary.
If you already run PostgreSQL, pgvector lets the vector live in the same row as the business data — one transaction, one backup, joins included. That convenience is a big reason to start there. See pgvector in practice.