Vector Database Handbook from zero to production Bipin Singh
Using a vector database

The data model

2 min readChapter 10 of 25By Bipin Singh

Different vector databases use different names, but almost all share the same basic model. Once you know it, switching between products is mostly a matter of vocabulary.

The building blocks

Concept Also called What it is
Collection Index, class, table A group of records that share a vector size and similarity metric
Record Point, object, row, document One stored item
ID Key, primary key A unique identifier you choose (or the database generates)
Vector Embedding The list of numbers used for similarity search
Metadata Payload, properties, attributes, columns Extra fields: text, tags, dates, numbers, IDs
Namespace Partition, tenant, shard key A sub-division inside a collection, often per customer

A single record might look like this:

{
  "id": "kb-1042#chunk-3",
  "vector": [0.012, -0.044, 0.318, "... 1,533 more numbers ..."],
  "metadata": {
    "doc_id": "kb-1042",
    "title": "Why was my payment refused?",
    "section": "Card payments",
    "language": "en",
    "updated_at": "2026-08-14",
    "tenant_id": "acme",
    "source_url": "https://example.com/help/payments/refused",
    "text": "If your card is declined at a shop, the most common reasons are..."
  }
}

Designing records well

Choose stable IDs. Derive IDs from your source data — document ID + chunk number — so re-running ingestion updates records instead of creating duplicates.

Store what you filter on. Any field you'll want to filter by (tenant, language, date, category, access group) must be in the metadata — and usually indexed. See Metadata filtering.

Store what you display, or a pointer to it. Search results need something to show. Either keep the text snippet in metadata, or keep an ID and fetch details from your main database. Keeping text in the vector database is simpler; keeping it elsewhere avoids duplication and size limits.

Record provenance. Keep the source document ID, the chunk position, and the embedding model name/version. You'll need them for updates, deletes, citations and re-embedding.

Key idea

The vector finds the record; the metadata makes the result useful — for filtering, display, permissions and maintenance.

One collection or several?

Use separate collections when… Use one collection with metadata when…
Vectors come from different embedding models or have different sizes Same model and size
Data has completely different access patterns or lifecycles Items are searched together
You need hard isolation (e.g. regulated customers) Logical separation by filter is enough

Multiple vectors per item

Sometimes one item deserves several vectors — a product with a text description and a photo, or a document with a title vector and a body vector. Options:

Where the source of truth lives

Treat the vector database as a search index, not your system of record. Keep original documents and business data in your primary database or storage, and design ingestion so the vector index can always be rebuilt from them. This makes model changes, migrations and recovery far less scary.

Tip

If you already run PostgreSQL, pgvector lets the vector live in the same row as the business data — one transaction, one backup, joins included. That convenience is a big reason to start there. See pgvector in practice.

Bipin Singh
Written by Bipin Singh

Senior Full-Stack Engineer · AI & AWS. I build production search, RAG and AI systems on AWS and Postgres.

Work with me