← All parts of this equation

Equation 1 · Part 2 · Embeddings and the Geometry of Similarity

Symbol v

sim⁡(u,v)=⟨u,v⟩∥u∥ ∥v∥.\operatorname{sim}(u, v) = \frac{\langle u, v \rangle}{\lVert u \rVert \, \lVert v \rVert}.
vv

What this part means

v is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Its job in the formula

v is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

The passage around this formula

The conventional similarity metric is the cosine of the angle between two vectors: sim⁡(u,v)=⟨u,v⟩∥u∥ ∥v∥\operatorname{sim}(u, v) = \frac{\langle u, v \rangle}{\lVert u \rVert \, \lVert v \rVert}. Cosine became conventional for reasons that are mostly good. It is scale-invariant, which matters when vector norms correlate with nuisance properties like token frequency or document length. It reduces to an inner product on normalised vectors, which is cheap and which most approximate indexes support natively. And it is what several influential embedding models were explicitly trained to make meaningful: Sentence-BERT fine-tuned siamese networks precisely so that sentence embeddings could be compared with cosine similarity, cutting a pairwise-comparison workload from roughly 65 hours to…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.