Robots Atlas>ROBOTS ATLAS
Retrieval

Cosine Similarity

ActivePublished
Key innovation
Measures the similarity of vector direction while ignoring magnitude, making documents of different sizes comparable.
Category
Retrieval
Abstraction level
Primitive
Operation level
RetrievalData
Use cases
Document ranking in search enginesSemantic search over embeddingsRecommendation systemsText deduplication and clusteringNearest-neighbour search (k-NN)

How it works

cos(θ) = (A · B) / (||A|| · ||B||). The dot product of vectors A and B divided by the product of their Euclidean norms. Result in [-1, 1] (for non-negative vectors such as TF-IDF, in [0, 1]): 1 = identical direction, 0 = orthogonality (no shared features).

Problem solved

Euclidean distance between document vectors is dominated by their length — a longer document is "farther" despite identical topic. Cosine similarity normalises this by looking only at the angle.

Implementation

Implementation pitfalls
Numerical instability with zero vectorsMedium

A zero vector (e.g. a document with no known terms) causes division by zero in the norm.

Fix:Add an epsilon to the denominator or filter out empty vectors before comparison.
Skipping L2 normalisation for repeated queriesLow

Recomputing norms for the same vectors repeatedly wastes time.

Fix:Pre-normalise vectors to unit length — then cosine reduces to a plain dot product.

Hyperparameters (configurable axes)

L2 pre-normalisationMedium

Whether vectors are pre-normalised to unit length — then cosine = dot product.

trueStandard for repeated queries against the same index.

Computational complexity

Time complexity: O(d) na parę, O(n·d) batch. Space complexity: O(d) na wektor.

Execution paradigm

Primary mode
Dense
Activation pattern
All paths active

Parallelism

Parallelism level
Fully parallel
Scope
Inference

Hardware requirements

Good fit

Batched cosine similarity is matrix multiplication — ideal for GPUs with dense embeddings.

Good fit

For sparse vectors (TF-IDF) a SIMD-capable CPU is efficient.