← All parts of this equation

Equation 13 · Part 7 · How Multimodal Models Actually Handle Video, Audio, and Space

Symbol d

C(r)=∫tntfT(t) σ(r(t)) c(r(t),d) dt,T(t)=exp⁡ ⁣(−∫tntσ(r(s)) ds),C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\, \sigma(\mathbf{r}(t))\, \mathbf{c}(\mathbf{r}(t), \mathbf{d})\, dt, \qquad T(t) = \exp\!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right),
dd

What this part means

d is one of the signed contributions combined to compute the quantity on the left.

Its job in the formula

d is one of the signed contributions combined to compute the quantity on the left.

The passage around this formula

The first treats a scene as a continuous field rather than a discrete grid at all. Neural radiance fields represent a scene as a fully connected network mapping a continuous 5D coordinate — a 3D position plus a 2D viewing direction — to a volume density and a view-dependent emitted colour, then use classical volume rendering to synthesize the colour a camera ray would see by integrating along it [ 8 ] . The rendering equation itself is the cleanest statement of what “continuous” buys and costs: C(r)=∫tntfT(t) σ(r(t)) c(r(t),d) dt,T(t)=exp⁡ ⁣(−∫tntσ(r(s)) ds)C(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\, \sigma(\mathbf{r}(t))\, \mathbf{c}(\mathbf{r}(t), \mathbf{d})\, dt, \qquad T(t) = \exp\!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right). where σ\sigma is volume density, c\mathbf{c} is emitted colour, and T(t) is accumulated transmittance along the ray up to t . There is no patch, no voxel grid, no fixed token count…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.