Equation 9 · The Hardest Unsolved Problems in AI Agent Architecture
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol b
the agent’s probability distribution over possible states before an observation.
Symbol s
everything true about the task — every file changed, every external call made, every fact established —.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Starting index or lower bound: s
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is . with a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…
Read the full surrounding passage
A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is . with a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the problem, and that reducing the reliability of a single observation channel in their own worked example forces a much larger plan graph, and a correspondingly larger memory requirement, just to stay confident [ 4 ] . Applied to a long-running LLM agent, s is everything true about the task — every file changed, every external call made, every fact established — and b is whatever the agent’s context window currently represents. The update above is exact only if the full state and its full history are kept; an agent cannot keep them, because the context window is finite and every token inside it costs money and attention.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to The Hardest Unsolved Problems in AI Agent Architecture