A trace is not an event
A seismograph never observes an earthquake. It observes one pier, in one vault, moving slightly. Everything anyone says afterwards about depth, magnitude, rupture direction and fault geometry is inferred from that trace and from a model connecting sources to traces. The instrument is honest about this: it is obviously a pendulum on a stone block, and nobody mistakes the scratched line for the fault.
Neural recording works the same way and is far less obvious about it. An electrode does not record thought, intention, memory or attention. It records a voltage difference between its tip and a reference, produced by ionic currents crossing membranes near it, filtered by the tissue in between. A microscope does not record spiking; it records photons emitted by an engineered protein whose brightness depends on calcium concentration, which depends on spiking through several intervening steps. A scanner does not record neural activity; it records a magnetic resonance signal whose contrast depends on blood oxygenation, which follows neural activity by seconds.
Between each of those physical measurements and any published claim about the brain there is a chain of inference. Each link is defensible. Each link also discards something specific, and the discarded thing is rarely stated in the sentence that reaches the reader. What follows is an attempt to make the chain explicit, modality by modality, and then to look at what the chain implies for the two arguments the field spends most of its energy on — causation, and replication.
What an extracellular electrode is actually in contact with
Start with the most direct measurement available. A metal or silicon electrode placed in the extracellular space senses the potential produced by transmembrane currents in the surrounding tissue. Under the standard volume-conductor treatment, with the medium approximated as homogeneous, isotropic and purely resistive, and each current source treated as a point, the potential at position
Three assumptions are visible in that expression, and all three are approximations rather than facts. The medium is not homogeneous. Conductivity may be frequency-dependent. And the sum runs over every current source in range, not over the neuron of interest.
The physiological content matters as much as the geometry. Buzsáki, Anastassiou and Koch, reviewing the origins of extracellular fields, are explicit that the dominant contributor to the extracellular signal is synaptic transmembrane current, with sodium and calcium spikes, fluxes through voltage- and ligand-gated channels, and intrinsic membrane oscillations all capable of substantially shaping the field [1]. An extracellular trace is therefore not a recording of output. It is a recording of a mixture in which input currents usually dominate.
The inverse-distance term does the other important work. Amplitude falls off with distance from the source, which is why a probe can resolve individual spikes at all — nearby cells produce large deflections, distant ones blur into the background. That same term is the origin of the sampling bias examined below.
Spike sorting is an inference step, not a preprocessing step
The transition from a continuous voltage trace to a list of neurons with firing times is called spike sorting, and it is the point at which a physical measurement becomes a statement about identified cells. High-pass filtering isolates fast deflections; a threshold detects candidate events; features are extracted from the waveform shapes; and a clustering step assigns events to putative single units. Rey, Pedreira and Quian Quiroga, in their review of the field’s history and prospects, describe this as a crucial step for extracting information from extracellular recordings and note that new electrodes monitoring hundreds of neurons simultaneously make the algorithmic problem both more exciting and more challenging [4].
The ambiguity is not hypothetical, and it has been measured against ground truth. Harris, Henze, Csicsvari, Hirase and Buzsáki compared tetrode spike separation with simultaneous intracellular recording from one of the sorted cells, which gives an unarguable answer for that cell. They reported that manual clustering produced errors typically ranging from zero to thirty per cent, depending on spike amplitude, firing pattern, the similarity of neighbouring waveforms and the experience of the operator, while their semi-automatic classification reduced errors to a range of zero to eight per cent [3]. Two facts follow. Sorting error is real and non-trivial. And it depends on who did the sorting — a property no physical measurement should have.
Two failure modes deserve naming because they push results in opposite directions. Over-splitting divides one neuron’s spikes across several clusters, inflating the count of recorded units and lowering each one’s apparent firing rate. Over-merging pools several neurons into one cluster, producing an artificial unit with a mixed tuning profile that can look like a cell with surprisingly complex selectivity. Bursting cells, whose spikes shrink in amplitude within a burst, are systematically vulnerable to both. When a paper reports that a population of neurons responded to a stimulus, the population is a product of these decisions.
The probe chooses its neurons, and it does not choose fairly
High-density silicon probes have changed the scale of the problem without changing its structure. The Neuropixels design described by Jun, Steinmetz, Siegle and colleagues integrates 960 recording sites along a single shank with 384 simultaneously addressable channels; in one reported experiment two probes isolated 741 putative single neurons across five brain structures, with single probes typically isolating twenty to two hundred neurons per structure [5]. Those are large numbers relative to the few dozen units per shank that earlier technology yielded. They are very small numbers relative to the population they sit inside.
More importantly, the subset is not random. Because amplitude falls with distance, a probe preferentially detects cells that are close to it and cells that are large, since larger somata and thicker dendrites carry larger transmembrane currents. Cells that fire rarely may never cross the detection threshold at all. Shoham, O’Connor and Segev pressed this point directly, arguing that many brain areas are far more sparsely active than commonly assumed and that a population of neurons firing rarely or only to highly specific stimuli — which they called dark neurons, by analogy with the astrophysical case — would be systematically under-represented by standard extracellular methods [6].
This is an argument about the denominator, and it is worth separating fact from interpretation. The fact is that detectability depends on amplitude and rate. The interpretation, which remains contested, is how much of the population that biases away. A recording that finds thirty per cent of sampled cells tuned to a variable has not established that thirty per cent of the local population is tuned to it; it has established that thirty per cent of the detectable, sortable, sufficiently active subset was.
The local field potential is neither reliably local nor purely a field
Low-pass filtering the same electrode signal yields the local field potential, conventionally read as an index of synaptic input and population synchrony in the neighbourhood of the tip. The word local is doing a great deal of work, and whether it is earned is a live dispute.
Kajikawa and Schroeder compared local field potential, current source density and multiunit activity in macaque auditory cortex and concluded that the three signals have systematically different listening areas, ordered multiunit activity, then current source density, then field potential. They reported direct evidence of passive spread of field potentials to sites more than a centimetre from their origins, and described the recorded signal as a mixture of local potentials with volume-conducted potentials from distant sites [2]. Modelling work in the same period, reviewed by Buzsáki and colleagues, emphasises that the effective spatial reach depends strongly on the geometry and correlation structure of the underlying sources rather than being a fixed property of the tissue [1].
The disagreement is not about the physics, which both sides accept. It is about which regime the cortex is usually in. Synchronous input to an aligned population of pyramidal cells produces a far-reaching dipole field; asynchronous input to a less ordered population does not. Both camps agree that a field potential recorded during a strongly synchronised state should be treated as a regional signal, and that attributing it to a specific column requires an argument, not a default.
Calcium indicators are a slow chemical proxy for a fast electrical event
Optical imaging with genetically encoded calcium indicators trades directness for population coverage and cell-type specificity. The measured quantity is fluorescence, which reports calcium binding, which reports calcium influx, which follows action potentials. Each arrow is a physical step with its own time constant, and the composite is slow relative to spiking.
Chen, Wardill, Sun and colleagues, introducing the GCaMP6 family, reported that under favourable conditions individual spikes could be resolved when separated by roughly 100 to 150 milliseconds for GCaMP6s, 75 to 100 for GCaMP6m and 50 to 75 for GCaMP6f, with GCaMP6f showing a rise time about twofold faster and a decay about 1.7-fold faster than GCaMP5G; they reported single-action-potential detection at 99 per cent, plus or minus 0.2, at a one per cent false-positive rate in pyramidal neurons under their conditions [7]. That is a developer’s characterisation of a new tool under conditions chosen to show what it can do, and should be read as such.
The relevant question for anyone interpreting a published imaging result is what happens under the conditions people actually image in. Huang, Knoblich, Ledochowitsch and colleagues addressed this with simultaneous loose-patch electrophysiology and two-photon imaging in GCaMP6 transgenic mice. Under optimal high-resolution imaging they reported single-spike detection probability of 0.70, plus or minus 0.06, in eighteen Emx1-s neurons and 0.40, plus or minus 0.08, in nine Cux2-f neurons at a one per cent false-positive rate, with two-spike events detected far more reliably. When the images were downsampled to approximate the conditions of population imaging across a wide field of view, single-spike detection for Emx1-s neurons fell to 0.32, plus or minus 0.05 [8].
These two results are not in contradiction; they characterise different regimes, and expression system, indicator variant, imaging rate and field of view all differ. Taken together they support a specific reading: population calcium imaging is a good detector of periods of elevated firing and a poor detector of isolated single spikes. Any analysis that treats a deconvolved fluorescence trace as a spike train inherits that asymmetry, and results that depend on precise single-spike timing are the ones most exposed.
EEG sums synchronous currents and then loses their address
Scalp electroencephalography is the oldest non-invasive method and the one whose forward model is best understood in principle. The standard account holds that simultaneous postsynaptic potentials across a population of similarly oriented neurons produce an extracellular dipole field that superposes and propagates through brain, cerebrospinal fluid, skull and scalp to the electrode.
Two consequences follow from that description alone. First, EEG is selective for synchrony and for geometry: activity that is intense but asynchronous, or arises in a population without consistent dendritic orientation, can be almost invisible at the scalp while a weaker but well-aligned and well-synchronised source dominates. An absence in the EEG is therefore weak evidence of an absence in the brain. Second, recovering source locations from scalp voltages is an inverse problem with no unique solution, so every source estimate depends on regularisation choices and a head model.
Cohen, surveying the field, argues that despite the method’s centrality we know remarkably little about how neural circuit activity produces the specific EEG features that get linked to cognition, and that closing this gap requires connecting scalp measurements to circuit-level neuroscience rather than treating the standard model as settled [9]. This is an argument about the maturity of the forward model, not a rejection of the method, and it is a useful corrective to the frequency-band vocabulary in which EEG findings are often reported as though the bands were mechanisms.
BOLD is a haemodynamic signal, and the gap is the interesting part
Functional MRI’s blood-oxygen-level-dependent contrast is the modality furthest from spiking, and the one whose interpretation has been argued most explicitly by its own practitioners. The measurement responds to changes in deoxyhaemoglobin concentration produced by local changes in blood flow, volume and oxygen consumption. The haemodynamic response peaks several seconds after the neural event that triggers it.
Nearly all standard analysis assumes that neural activity and the measured signal are related by a linear time-invariant system, so that the measured time course is a convolution of an underlying activity time course with a canonical response function plus noise:
Writing it this way exposes what is being assumed rather than measured: that
The empirical question of what
The practical translation is narrow and worth stating plainly. An activation is a regional change in a slow vascular signal that best tracks synaptic input. It does not distinguish excitation from inhibition, since inhibitory synaptic activity is metabolically expensive too. It does not identify which cells were involved. And a region that is more active is not thereby the region that performs the function.
What each modality buys and what it pays
The trade-offs are structural rather than technological, and it helps to state them as such. Intracellular recording gives the membrane potential of one cell with sub-millisecond precision and a sample size of one. Extracellular arrays give millisecond spike times from tens to hundreds of biased-sample cells in one or a few locations. Calcium imaging gives thousands of identified cells with cell-type specificity, at a temporal resolution set by indicator kinetics rather than by the camera. EEG and magnetoencephalography give millisecond whole-head coverage with source locations that are estimates. Functional MRI gives whole-brain coverage at millimetre spatial resolution and a temporal resolution set by vascular physiology.
No method occupies the top-left corner of coverage against fidelity, and the arrangement is not an accident of engineering. Sampling more cells means either smaller signals per cell or a slower reporter; sampling non-invasively means summing over sources; summing over sources means losing the address. Claims that a new instrument has eliminated the trade-off — rather than moved along it — are vendor assertions until an independent group demonstrates the same numbers under ordinary experimental conditions.
From correlation to cause, and what interventions really add
Every measurement discussed so far is observational. A neuron whose firing correlates with a decision may be causing the decision, reporting it, receiving a copy of it, or tracking something else that correlates with all of the above. This is the point at which perturbation methods enter, and their contribution is real.
Optogenetics made the perturbation fast and targeted. Boyden, Zhang, Bamberg, Nagel and Deisseroth showed that channelrhodopsin-2, delivered genetically and driven optically, allowed control of neural firing and synaptic transmission on a millisecond timescale in defined populations [12]. Combined with cell-type-specific promoters, this converts a correlational claim into a testable causal one: if this population drives the behaviour, driving it should change the behaviour.
The inference from a perturbation to a function is nevertheless not automatic, and the sharpest published warning comes from within the method’s own community. Otchy, Wolff, Rhee and colleagues found that transient inactivation of motor cortex in rats and of nucleus interface in songbirds severely degraded learned task-specific movement patterns and courtship song — behaviours that recover spontaneously after permanent lesions of the very same areas [13]. The acute manipulation and the chronic lesion give opposite answers about whether the area is required. The authors attribute the discrepancy to acute off-target effects on downstream circuits that are not present once the system has had time to compensate.
The methodological lesson generalises beyond those two systems. An acute perturbation tests whether ongoing activity in a region is currently participating in a behaviour. A chronic lesion tests whether the system can eventually perform the behaviour without that region. These are different questions, and both are different again from asking what the region computes. Krakauer, Ghazanfar, Gomez-Marin, MacIver and Poeppel push the argument one level further, contending that detailed analysis of tasks and the behaviour they elicit is what identifies the component processes and algorithms needing explanation, and that neural intervention tests causality against those processes rather than substituting for defining them [20]. Their claim that behavioural decomposition should generally precede neural investigation is a position in an ongoing methodological argument, not a consensus.
The statistics behind the replication debate
The measurement chain sets an upper bound on what a dataset can support. Statistical practice determines how much of that bound is actually used, and three specific issues account for much of the published concern.
The first is non-independence. Kriegeskorte, Simmons, Bellgowan and Baker named the practice of using the same data both to select a region or set of voxels and to perform the selective analysis on it, showing that this yields distorted descriptive statistics and invalid inference whenever the result statistic is not independent of the selection criterion under the null [14]. Their accompanying survey of fMRI papers published in 2008 in five leading journals reported that a substantial fraction — 42 per cent of 134 papers — performed at least one non-independent selective analysis. The remedy is structural rather than statistical: define regions on independent data, or on anatomy, or cross-validate.
The second is multiple comparisons in a spatially correlated field. Eklund, Nichols and Knutsson used resting-state data from 499 healthy controls as null data to run three million task group analyses across the standard software packages. They found the common parametric methods conservative for voxelwise inference but invalid for clusterwise inference, with false-positive rates reported as high as around 70 per cent for the most permissive cluster-defining thresholds, and found that nonparametric permutation testing produced nominal rates for both [16]. The paper’s scope has itself been argued over in subsequent exchanges in the same journal, and the disagreement concerns how many published results are actually affected, not whether the parametric cluster assumption can fail.
The third is analytic flexibility. Botvinik-Nezer, Holzmeister, Camerer and colleagues gave one fMRI dataset and nine pre-specified hypotheses to 70 independent teams. No two teams chose identical workflows, and the resulting hypothesis-test outcomes varied substantially, even between teams whose intermediate statistical maps were highly correlated; a meta-analytic aggregation across teams nevertheless recovered a consensus set of activated regions [17]. That last clause matters. The finding is not that the data were meaningless but that the pipeline is a variable, and an unreported one.
Underneath all three sits statistical power. Button, Ioannidis, Mokrysz and colleagues made the argument that low power not only reduces the chance of detecting a true effect but also reduces the probability that a statistically significant result reflects one, while inflating the effect sizes of those results that do reach significance [15]. In a field where a subject is expensive and an implanted animal more so, small samples are an economic fact, which makes the inflation systematic rather than occasional.
The variable no one records
The deepest limitation is not resolution and not statistics. It is that most theories in systems neuroscience are about a quantity that no instrument currently observes: the joint state of a defined circuit, at the resolution of individual synaptic weights and cell-type identities, in a behaving animal, over the timescale on which the behaviour unfolds. Every method described here supplies a projection of that state — a biased subsample, a spatial average, a temporal blur, or a vascular echo.
Two results sharpen why the gap is not merely quantitative. Marder and Goaillard, reviewing variability and homeostasis in neurons and networks, showed that similar circuit output can arise from substantially different underlying parameter sets, so that measuring the output does not identify the parameters that produced it [19]. Degeneracy of this kind means the inverse problem is not just noisy but genuinely many-to-one. And Jonas and Kording applied a battery of standard neuroscience analyses to a system whose ground truth is fully known — the MOS 6502 microprocessor — and reported that the approaches revealed interesting structure while failing to describe the hierarchy of information processing, concluding that current analytic approaches may fall short of producing meaningful understanding regardless of the amount of data [18]. Their study is a demonstration on one artificial system and its transfer to biological circuits is debated; the demonstration nevertheless stands as the field’s most explicit test of its own analysis toolkit against a known answer.
What would change this, and how it could be shown wrong
A prediction, with its terms stated. Horizon: the next five years, to 2031. Assumptions: continued growth in probe channel counts and in indicator speed, no fundamental change in the physics of any modality, and continued expansion of standardised open datasets.
The prediction is that the binding constraint on systems neuroscience over that horizon will shift further from data volume toward the validity of the inference chain, and that the visible response will be procedural rather than instrumental. Observable indicators would include ground-truth validation datasets — simultaneous intracellular and extracellular, or paired ephys and imaging — becoming an expected component of methods papers rather than a specialist genre; multi-pipeline reporting appearing in primary imaging papers; and spike-sorting uncertainty being propagated into downstream population analyses instead of being resolved before them.
The disconfirmation condition is specific. If, by 2031, several independent groups demonstrate a non-invasive method that resolves identified single neurons at millisecond precision across a behaving mammalian brain — or if large-scale reanalyses show that conclusions from spike-sorted, cluster-thresholded or deconvolved data are robust to the full range of defensible processing choices — then the constraint was data volume after all, and this prediction was wrong.
The seismological analogy holds to the end, including in what it permits. Seismology built an accurate picture of the Earth’s interior from instruments that only ever measured the motion of their own piers. It did so by characterising the instrument precisely, modelling the propagation path explicitly, deploying many stations rather than one, and stating what a trace could and could not distinguish. None of that required observing the source directly. It required refusing to confuse the trace with the event.