← Back to article

Equation 9 · The Hardest Unsolved Problems in AI Agent Architecture

What does this equation mean?

b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s),b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationseta O(o mid s', a) sum_s T(s' mid s, a) b(s)
Result or conditionb'(s')
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

bb

Symbol b

the agent’s probability distribution over possible states before an observation.

Understand this part →

ss

Symbol s

everything true about the task — every file changed, every external call made, every fact established —.

Understand this part →

η\eta

Symbol eta

the normalizing constant.

Understand this part →

OO

Symbol O

the observation model.

Understand this part →

oo

Symbol o

the observation that followed.

Understand this part →

aa

Symbol a

the action just taken.

Understand this part →

TT

Symbol T

the environment’s transition model.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

ss

Starting index or lower bound: s

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s)b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s). with η\eta a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…
Read the full surrounding passage
A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s)b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s). with η\eta a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the problem, and that reducing the reliability of a single observation channel in their own worked example forces a much larger plan graph, and a correspondingly larger memory requirement, just to stay confident [ 4 ] . Applied to a long-running LLM agent, s is everything true about the task — every file changed, every external call made, every fact established — and b is whatever the agent’s context window currently represents. The update above is exact only if the full state and its full history are kept; an agent cannot keep them, because the context window is finite and every token inside it costs money and attention.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to The Hardest Unsolved Problems in AI Agent Architecture

See this formula across 1 published context →

Browse the mathematical compendium →