← All parts of this equation

Equation 4 · Part 1 · Mechanistic Interpretability in 2035: Scenarios and Falsifiers

Symbol R

R(t)=1 ⁣[A(t)≥a∗]⋅1 ⁣[S(t)≥s∗]R(t) = \mathbb{1}\!\left[A(t) \ge a^{*}\right] \cdot \mathbb{1}\!\left[S(t) \ge s^{*}\right]
RR

What this part means

R is part of the quantity the equation computes from the expression on the right.

Its job in the formula

R is part of the quantity the equation computes from the expression on the right.

The passage around this formula

…result covering a hand-picked sliver of the model fail for different reasons. Admissibility is therefore better modelled as a conjunction of two thresholds than a weighted sum of two moving averages: R(t)=1 ⁣[A(t)≥a∗]⋅1 ⁣[S(t)≥s∗]R(t) = \mathbb{1}\!\left[A(t) \ge a^{*}\right] \cdot \mathbb{1}\!\left[S(t) \ge s^{*}\right]. R(t) stays at zero however high either term climbs alone, exactly as the AISI account above already anticipates by asking for outside validation before treating interpretability-based detection as sufficient on its own [ 11 ] . Axis A determines whether A(t) can plausibly clear a∗a^{*} within the…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.