Equation 1 · Embeddings and the Geometry of Similarity
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol u
u is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol v
v is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →Denominator: lVert u rVert lVert v rVert
The complete quantity below the fraction bar; it must be nonzero for this division.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
The conventional similarity metric is the cosine of the angle between two vectors: . Cosine became conventional for reasons that are mostly good. It is scale-invariant, which matters when vector norms correlate with nuisance properties like token frequency or document length. It reduces to an inner product on normalised vectors, which is cheap and which most approximate indexes support natively. And it is what several influential embedding models were explicitly trained to make meaningful: Sentence-BERT fine-tuned siamese networks precisely so that sentence embeddings could be compared with cosine similarity, cutting a pairwise-comparison workload from roughly 65 hours to…
Read the full surrounding passage
The conventional similarity metric is the cosine of the angle between two vectors: . Cosine became conventional for reasons that are mostly good. It is scale-invariant, which matters when vector norms correlate with nuisance properties like token frequency or document length. It reduces to an inner product on normalised vectors, which is cheap and which most approximate indexes support natively. And it is what several influential embedding models were explicitly trained to make meaningful: Sentence-BERT fine-tuned siamese networks precisely so that sentence embeddings could be compared with cosine similarity, cutting a pairwise-comparison workload from roughly 65 hours to about 5 seconds while preserving accuracy [ 5 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to Embeddings and the Geometry of Similarity