← All parts of this equation

Equation 5 · Part 7 · From n-Grams to Reasoning Models: A Technical History of the Language Model

Symbol k

ci=∑j=1Txαijhj,αij=exp⁡(eij)∑k=1Txexp⁡(eik).c_i = \sum_{j=1}^{T_x} \alpha_{ij} h_j, \qquad \alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{k=1}^{T_x} \exp(e_{ik})}.
kk

What this part means

k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Its job in the formula

k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

The passage around this formula

Bahdanau, Cho and Bengio proposed letting the decoder search the source for the parts relevant to each output word, rather than reading from a single compressed vector [ 9 ] . The decoder computes, at each output step i , a context vector as a weighted sum of all encoder states: ci=∑j=1Txαijhj,αij=exp⁡(eij)∑k=1Txexp⁡(eik)c_i = \sum_{j=1}^{T_x} \alpha_{ij} h_j, \qquad \alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{k=1}^{T_x} \exp(e_{ik})}. The capacity of the intermediate representation now grows with the input rather than being fixed in advance. They further reported that the learned alignments corresponded well with human linguistic intuition — an interpretability result that arrived free with a performance fix, which is rare.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.