← Back to article

Equation 8 · Structure from Sequence: What Protein Folding Prediction Did and Did Not Settle

What does this equation mean?

P(a1,…,aL)=1Zexp⁡ ⁣[∑ihi(ai)+∑i<jeij(ai,aj)],P(a_1,\ldots,a_L) = \frac{1}{Z}\exp\!\left[\sum_{i} h_i(a_i) + \sum_{i<j} e_{ij}(a_i,a_j)\right],

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start with1
Divide byZ
This relates toP(a_1,ldots,a_L)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

PP

Symbol P

P is part of the quantity the equation computes from the expression on the right.

Understand this part →

a1a_1

Symbol a_1

a1a_1 is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

aLa_L

Symbol a_L

aLa_L is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

ZZ

Symbol Z

Z occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

ii

Symbol i

i appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

hih_i

Symbol h_i

the single-site fields.

Understand this part →

aia_i

Symbol a_i

aia_i is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

jj

Symbol j

j appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

eije_{ij}

Symbol e_ij

eie_ij is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

aja_j

Symbol a_j

aja_j is one of the signed contributions combined to compute the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
addition

addition

Add the term after the plus sign to the term or group before it.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

11

Numerator: 1

The complete quantity above the fraction bar.

Understand this part →

ii

Starting index or lower bound: i

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

i<ji<j

Starting index or lower bound: i<j

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

The fix is to fit a global model that explains the observed column statistics with the smallest number of direct couplings. In the maximum-entropy formulation used in this literature, the probability of a full sequence takes a Potts form, P(a1,…,aL)=1Zexp⁡ ⁣[∑ihi(ai)+∑i<jeij(ai,aj)]P(a_1,\ldots,a_L) = \frac{1}{Z}\exp\!\left[\sum_{i} h_i(a_i) + \sum_{i<j} e_{ij}(a_i,a_j)\right]. with single-site fields hih_i and pair couplings eije_{ij} fitted so that the model reproduces the single-column and pairwise frequencies of the alignment. Marks and colleagues reported that the strength of the inferred couplings is a strong predictor of residue proximity in the folded structure, and that the top-scoring couplings are accurate and well-distributed enough to define a three-dimensional fold [ 5 ] . Morcos and colleagues…
Read the full surrounding passage
The fix is to fit a global model that explains the observed column statistics with the smallest number of direct couplings. In the maximum-entropy formulation used in this literature, the probability of a full sequence takes a Potts form, P(a1,…,aL)=1Zexp⁡ ⁣[∑ihi(ai)+∑i<jeij(ai,aj)]P(a_1,\ldots,a_L) = \frac{1}{Z}\exp\!\left[\sum_{i} h_i(a_i) + \sum_{i<j} e_{ij}(a_i,a_j)\right]. with single-site fields hih_i and pair couplings eije_{ij} fitted so that the model reproduces the single-column and pairwise frequencies of the alignment. Marks and colleagues reported that the strength of the inferred couplings is a strong predictor of residue proximity in the folded structure, and that the top-scoring couplings are accurate and well-distributed enough to define a three-dimensional fold [ 5 ] . Morcos and colleagues developed a computationally efficient direct-coupling analysis that disentangled direct from indirect correlations and evaluated contact-prediction accuracy across a large number of protein domain families from sequence information alone [ 6 ] .

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Structure from Sequence: What Protein Folding Prediction Did and Did Not Settle

See this formula across 1 published context →

Browse the mathematical compendium →