Measuring similarity
"Close together" needs a precise definition before a computer can use it. Vector databases use a similarity metric (or its opposite, a distance metric) to score how alike two vectors are. Three metrics cover almost every real system.
We'll use three small 2-D vectors throughout:
A = [3, 4]B = [4, 3]C = [6, 8](same direction as A, twice as long)
Euclidean distance (L2)
The straight-line distance between the two points — what you'd measure with a ruler.
distance(A, B) = √((3−4)² + (4−3)²) = √(1 + 1) = √2 ≈ 1.41
distance(A, C) = √((3−6)² + (4−8)²) = √(9 + 16) = √25 = 5
Smaller means more similar. Notice that A and C are far apart even though they point the same way — Euclidean distance cares about length as well as direction.
Dot product (inner product)
Multiply matching positions and add them up.
A · B = 3×4 + 4×3 = 12 + 12 = 24
A · C = 3×6 + 4×8 = 18 + 32 = 50
Bigger means more similar. The dot product grows with both direction alignment and vector length, so long vectors score high. Some models are trained so that length carries meaning (for example, popularity in recommendations), and the dot product uses it.
Cosine similarity
Cosine similarity measures only the angle between vectors, ignoring their length. It is the dot product divided by both lengths:
cosine(A, B) = (A · B) / (|A| × |B|)
|A| = √(3² + 4²) = 5 |B| = 5 |C| = 10
cosine(A, B) = 24 / (5 × 5) = 0.96
cosine(A, C) = 50 / (5 × 10) = 1.00 ← identical direction
It ranges from −1 (opposite) through 0 (unrelated) to 1 (same direction). Cosine distance is simply 1 − cosine similarity.
Normalisation makes them agree
If you normalise vectors — scale each one to length 1 — then all three metrics rank results the same way, and the dot product equals cosine similarity. Many embedding models already return normalised vectors, and databases often normalise for you when you choose cosine.
function normalise(v: number[]): number[] {
const length = Math.sqrt(v.reduce((sum, x) => sum + x * x, 0));
return v.map((x) => x / length);
}
// normalise([3, 4]) → [0.6, 0.8]
With normalised vectors, prefer the dot product in your database: it gives the same ranking as cosine and is slightly cheaper to compute.
Which metric should you use?
| Situation | Use |
|---|---|
| Your embedding model's documentation recommends one | That one — always check first |
| Text embeddings (most common case) | Cosine, or dot product on normalised vectors |
| Vector length carries meaning (some recommendation models) | Dot product |
| Raw numeric features, physical measurements, some image features | Euclidean |
Choose the metric when you create the index. Building an index for Euclidean distance and querying it as if it were cosine gives silently wrong rankings.
Similarity is just arithmetic on lists of numbers. The hard part — covered next — is doing that arithmetic against millions of vectors in a few milliseconds.