Metadata filtering
Real searches almost always have conditions: only this customer's documents, only English, only products in stock, only things this user is allowed to see. Combining similarity search with filters is one of the hardest parts of vector database engineering — and one of the most important for correctness and security.
Three ways to filter
Post-filtering
Run the vector search, get the top results, then drop the ones that fail the filter.
search top 10 → filter → maybe only 2 left (or 0)
Simple, but if the filter is strict, you can end up with too few results — or none — even though matching items exist further down the ranking. Asking for more results (top 100, then filter) helps but wastes work and still isn't guaranteed.
Pre-filtering
Apply the filter first, then search only among matching vectors.
filter → 3,000 matching items → search those → top 10
Always returns enough correct results. But an ANN index is built over all vectors, so searching just a filtered subset efficiently needs special support. If the subset is small, brute force over it is fine; if it's large, the database needs a smarter approach.
Filtered search inside the index
Modern vector databases integrate the filter into the index traversal: while walking the HNSW graph, they skip nodes that don't match but still use them for navigation, or they keep extra links so filtered subsets stay connected. Many also switch strategies automatically — brute force when the filter matches few items, index search when it matches many.
Selectivity is what matters
Selectivity is the fraction of records that pass the filter.
| Filter matches | What tends to work best |
|---|---|
Most records (e.g. language = "en" on an English-heavy dataset) |
Index search with in-traversal filtering, or post-filtering |
| A medium fraction | In-traversal filtering with payload indexes |
| Very few records (e.g. one small tenant) | Pre-filter, then brute force over the subset |
A graph search that skips most nodes can wander into dead ends and return fewer or worse results. If filtered queries show low recall, check how your database handles restrictive filters — and test them explicitly.
Index your metadata fields
Just as in a relational database, filtering on a field is fast only if that field is indexed. Most vector databases let you create payload / metadata indexes on the fields you filter by — keyword, integer, date or boolean. Without them, filters may require scanning records.
Permissions are filters you must never get wrong
For enterprise search and RAG, access control is usually implemented as a metadata filter: each chunk stores which groups or users may see it, and every query filters on the current user's groups.
{ "filter": { "tenant_id": "acme", "allowed_groups": { "any": ["finance", "all-staff"] } } }
Apply permission filters inside the search, never after generating an answer. If unauthorised chunks reach an LLM's context, it can leak them even if you hide the source links.
Practical tips
- Keep filter fields simple and typed — strings, numbers, dates, booleans. Avoid filtering on large free-text fields.
- Test filtered recall separately from unfiltered recall; they can differ a lot.
- Watch for empty results and show a sensible fallback rather than nothing.
- Partition by tenant if most queries filter by tenant — see Multi-tenancy.