Ingestion, updates and deletes
Getting vectors into a database is easy for a demo and surprisingly subtle in production. Data changes, documents get deleted, embedding models get upgraded — and the index must keep up without serving stale or forbidden results.
The ingestion pipeline
source data → extract text → split into chunks → embed → upsert into vector DB
↑ |
└──────────────── track what changed ←──────────────────┘
- Extract text (and metadata) from the source: database rows, PDFs, web pages, tickets.
- Chunk long content into retrievable pieces.
- Embed each chunk with your embedding model.
- Upsert — insert, or update if the ID exists — into the vector database.
Batch everything
Both embedding APIs and vector databases are much faster with batches than with one item at a time.
const BATCH = 100;
for (let i = 0; i < chunks.length; i += BATCH) {
const batch = chunks.slice(i, i + BATCH);
const vectors = await embedMany(batch.map((c) => c.text)); // one API call per batch
await vectorDb.upsert(
batch.map((c, j) => ({ id: c.id, vector: vectors[j], metadata: c.metadata }))
);
}
Add retries with backoff for rate limits, and make the job resumable — log the last successful batch so a failure halfway through doesn't mean starting over.
Keeping data fresh
| Strategy | How it works | Good for |
|---|---|---|
| Full rebuild | Re-ingest everything on a schedule | Small datasets, simple setups |
| Incremental by timestamp | Re-process items whose updated_at changed since the last run |
Most business data |
| Change data capture / events | React to inserts, updates and deletes as they happen | Near-real-time freshness |
| Content hashing | Hash each chunk; only re-embed when the hash changes | Saving embedding cost on large corpora |
Hash the chunk text and store the hash in metadata. On re-ingestion, skip any chunk whose hash hasn't changed — it can cut embedding costs dramatically.
Updating a document
When a document changes, its chunks may change in number as well as content. The safe pattern:
- Re-chunk and re-embed the new version.
- Upsert the new chunks.
- Delete chunks from the old version that no longer exist (e.g. delete where
doc_id = Xandversion < new_version).
Forgetting step 3 leaves orphaned chunks that keep appearing in search results.
Why deletes are tricky
Many indexes — HNSW in particular — can't cheaply remove a node, because other nodes route through it. So databases usually mark records as deleted (tombstones) and filter them out of results, then clean up or rebuild in the background (compaction).
Practical consequences:
- Heavy delete or update churn can slowly degrade recall and waste memory until compaction runs.
- Deleted data may physically remain for a while — important for privacy requests. Check how and when your database purges it.
For "right to be forgotten" and similar requests, confirm that deletes are physically completed (and removed from backups per your policy), not just hidden from results.
Consistency: when can you read what you wrote?
Some vector databases make new writes searchable immediately; others index asynchronously, so a just-inserted record may not appear for a short time (eventual consistency). If your application needs read-after-write — for example, a user uploads a document and immediately searches it — check the database's consistency options or wait for the write to be acknowledged as indexed.
Changing the embedding model
Vectors from different models are not comparable, so a model change means re-embedding everything. Do it safely:
- Create a new collection (or a new named vector) for the new model.
- Backfill it in the background.
- Run your evaluation set against both and compare.
- Switch queries to the new collection; keep the old one briefly for rollback.
Design ingestion so the whole index can be rebuilt from source at any time. It turns model upgrades, schema changes and disasters into routine jobs.