← Back to article

Equation 1 · Mechanistic Interpretability in 2035: Scenarios and Falsifiers

What does this equation mean?

A(t)A(t)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

AA

Symbol A

A is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

tt

Symbol t

t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Whether interpretability evidence becomes admissible for a safety certification is not a third axis; it is what the other two jointly produce, and the joint requirement is a conjunction rather than an average. Write A(t) for the share of a frontier model’s decision-relevant behaviour with a validated, causally checked account — Circuit Tracing’s own figures are the best public anchor for where A(t) sits today [ 2 ] — and S(t) for the share of published interpretability results built on a method that has cleared an agreed, cross-lab benchmark rather than a proxy metric of the kind SAEBench found unreliable [ 10 ] . A regulator or a court asked to accept mechanistic evidence needs both a…
Read the full surrounding passage
Whether interpretability evidence becomes admissible for a safety certification is not a third axis; it is what the other two jointly produce, and the joint requirement is a conjunction rather than an average. Write A(t) for the share of a frontier model’s decision-relevant behaviour with a validated, causally checked account — Circuit Tracing’s own figures are the best public anchor for where A(t) sits today [ 2 ] — and S(t) for the share of published interpretability results built on a method that has cleared an agreed, cross-lab benchmark rather than a proxy metric of the kind SAEBench found unreliable [ 10 ] . A regulator or a court asked to accept mechanistic evidence needs both a guarantee about how much of the model the evidence covers and a guarantee that the method itself is not still under live dispute; a high-coverage result from a disputed method and a well-validated result covering a hand-picked sliver of the model fail for different reasons. Admissibility is therefore better modelled as a conjunction of two thresholds than a weighted sum of two moving averages:

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Mechanistic Interpretability in 2035: Scenarios and Falsifiers

See this formula across 5 published contexts →

Browse the mathematical compendium →