NICER turned a neutron star’s surface into a number you can calculate, not just observe

For forty years, finding a spectral line from a neutron star’s surface has meant hunting for a redshift. A transition of known laboratory energy, seen shifted in a star’s spectrum, hands back the compactness of the star that produced it — in principle a clean way to weigh an object no telescope can resolve. In practice the hunt has returned almost nothing usable, and the reason is not that neutron stars lack the physics to produce such lines. It is that a line was always being asked to do two jobs at once: measure the star and, if anyone wanted to push further, test the gravity that bends its light. A new paper, “A Neutron Star Is a Redshift Standard,” argues that a different measurement — NASA’s Neutron star Interior Composition Explorer, or NICER — has quietly split those two jobs apart, and that splitting them is what makes a confirmed line valuable for the first time.

NICER does not read a spectral line at all. It times the arrival of soft X-ray photons from the hot spots on a rotating millisecond pulsar’s surface to better than a hundred nanoseconds, and a Bayesian model of that rotating waveform — how the visible hot-spot area changes shape as light bends around the star, how the pulse brightens and dims with the star’s spin — returns the star’s mass and radius together, as a single measurement with no line involved anywhere in the chain [1, 2]. Feed that mass and radius into general relativity’s exterior solution and the star’s surface redshift falls out as an exact consequence of the geometry, known to a precision set entirely by how well NICER measured the waveform. A star’s surface redshift, in other words, has become something a mission can calculate rather than something a rare absorption feature has to reveal.

Read the full paper (PDF)

ADVERTISEMENT

The formula is exact once mass and radius are known, and two pulsars mark its range

The relation itself is old and simple. For a spherical, non-rotating exterior, the surface redshift zsz_s↗ follows from the mass MM↗ and radius RR↗ as

zs=(1−2GMRc2)−1/2−1 z_s = \left(1 - \frac{2GM}{Rc^2}\right)^{-1/2} - 1 ↗

— a function of the star’s compactness alone, so that whatever fraction of a percent of uncertainty NICER puts on MM↗ and RR↗ propagates directly into the redshift. The paper works through six NICER-mapped pulsars this way, and two of them mark the range worth remembering. PSR J0740+6620, the most massive precisely-timed neutron star known, comes in at roughly 2.07 M⊙2.07\,M_\odot↗ with a radio-timing mass fixed independently by Shapiro delay to 2.08±0.07 M⊙2.08\pm0.07\,M_\odot↗ [4], giving a compactness near 0.25 and a surface redshift zs≃0.40z_s\simeq0.40↗ — the deepest gravitational potential in the sample, though also, because its radius is the harder-measured half of the pair, the least precisely known redshift, at roughly fifteen to sixteen percent. PSR J0437-4715 sits at the opposite end: the nearest and brightest rotation-powered millisecond pulsar NICER has targeted, with a radio-timing mass prior tight enough to pin the pair at M=1.418±0.037 M⊙M=1.418\pm0.037\,M_\odot↗ and R=11.36−0.63+0.95R=11.36^{+0.95}_{-0.63}↗ km [3], which maps to zs=0.257±0.026z_s=0.257\pm0.026↗ — a ten-percent redshift standard, the tightest of the six, built entirely from timing and pulse-shape data with no spectroscopy anywhere in the derivation.

A control-room console monitor glowing with a soft-focus mass-radius posterior contour, a signal cable unplugged and dangling beside the console
Figure 4. This is the shape the whole argument turns on: a closed contour in mass and radius, like the one NICER now returns for PSR J0437-4715, that fixes a surface redshift to about ten percent before a single line photon is counted [@choudhury-2024].

For forty years the order of operations ran backward, and NICER reverses it

Before NICER, there was no way to get MM↗ and RR↗ for a bursting neutron star independent of a spectral line, so the line had to carry the whole inference: assume an identification, extract a redshift, and only then estimate the equation of state consistent with that redshift and some other mass constraint. That order of operations has a structural weakness the paper’s authors state plainly — a line asked to fix the star’s structure has nothing left over to test the physics of the redshift itself, because any discrepancy could always be waved away as an equation-of-state uncertainty rather than a genuine anomaly. NICER breaks that circularity by supplying zsz_s↗ from a channel that never touches a transition energy. Once the redshift is known first and independently, a credibly identified surface line stops being a way to weigh the star and becomes a way to test whether matter actually falls the way general relativity says it should.

That test would not be a marginal addition to existing equivalence-principle work. The emitting atoms on a neutron-star surface sit in a gravitational potential Φs/c2\Phi_s/c^2↗ of order 0.2 to 0.4 — a QCD-dominated binding environment seven to eight orders of magnitude deeper than anything a terrestrial free-fall experiment reaches. MICROSCOPE’s satellite test of the weak equivalence principle, the most precise ever flown, bounds a titanium-platinum pair’s differential free fall at η=(−1.5±2.3±1.5)×10−15\eta=(-1.5\pm2.3\pm1.5)\times10^{-15}↗ [13], but at a potential difference near 10−1010^{-10} — many orders shallower than a neutron star’s surface. In the dilaton-coupling framework Thibault Damour and John Donoghue laid out for exactly this kind of composition-dependent violation, the induced effect scales as the inverse cube root of atomic mass number, a specific enough prediction that a surface line and a laboratory test genuinely probe different combinations of the same underlying couplings rather than one simply outranking the other [11]. A neutron-star surface line, laid against a NICER redshift standard, would not replace MICROSCOPE. It would open an axis MICROSCOPE cannot reach at all.

The line that was supposed to prove this already: 2002’s redshift in EXO 0748-676

The closest anyone came to this measurement, decades before NICER existed, was Jean Cottam, Frits Paerels, and Mariano Mendez’s 2002 analysis of the bursting neutron star EXO 0748-676. Stacking twenty-eight thermonuclear X-ray bursts observed with the Chandra grating spectrometer, they reported narrow absorption features matching iron and oxygen transitions, all consistent with a single redshift of z=0.35z=0.35↗ — a clean, quantitative result, and for several years the field’s best evidence that a neutron-star surface line could actually be resolved [5]. The claim implied a star built from ordinary nuclear matter, in a mass range of roughly 1.3 to 2.0 solar masses, and it was treated for most of a decade as the field’s proof of concept.

ADVERTISEMENT

It did not survive. A deeper 2003 XMM-Newton observation of the same source, by the same authors, gathered roughly two orders of magnitude more burst counts and did not reproduce the original features at anything like the claimed strength — a non-detection reported by the very team that made the original claim. The more decisive blow came from an entirely different measurement: Duncan Galloway, Jinrong Lin, Deepto Chakrabarty, and Jacob Hartman’s 2010 discovery of a 552-hertz burst oscillation in EXO 0748-676’s X-ray timing data, fixing the star’s spin at roughly ten times faster than the forty-five-hertz upper limit that had been assumed when the original line claim was made [6]. A star spinning that fast smears any surface line by simple rotation, and Lin, Feryal Özel, Chakrabarty, and Dimitrios Psaltis showed in a dedicated follow-up that the smearing implied by the measured spin is quantitatively incompatible with the narrow line widths originally reported — not a subtle tension, but a direct contradiction between two measurements of the same star [7]. The line was not falsified by a flaw in the 2002 spectra. It was falsified by a later, independent measurement of the one quantity — spin — that the original analysis had to assume rather than measure.

Nested X-ray mirror-module foil segments held in a machined metal alignment fixture, one shell suspended a hair above its slot
Figure 3. Every mass and radius in the paper's table reaches Earth through optics built to this tolerance; the nested-shell geometry that focuses grazing-incidence X-rays is the unglamorous physical reason a millisecond pulsar's soft X-ray photons can be timed precisely enough to fix a redshift at all.

A measured spin killed the line, and the number is bigger than people assume

The mechanism behind that contradiction is simple enough to state in one line, and the paper’s authors compute it carefully because the commonly quoted version understates it by an order of magnitude. Surface rotation Doppler-broadens any spectral line by a fraction of its energy equal to the star’s equatorial velocity over the speed of light, ΔE/E≈veq/c=2πνR/c≈1.3×10−1(ν/500 Hz)(R/12 km)\Delta E/E \approx v_{eq}/c = 2\pi\nu R/c \approx 1.3\times10^{-1}(\nu/500\,\text{Hz})(R/12\,\text{km})↗, where ν\nu↗ is the star’s spin frequency and RR↗ its radius.

Evaluated at a fairly typical burster spin of 500 hertz and a 12-kilometre radius, that formula gives a fractional smearing of about thirteen percent — not the roughly one-percent figure often casually quoted for rotational broadening, but an order of magnitude larger. At EXO 0748-676’s actual measured spin of 552 hertz, the smearing works out to roughly fourteen percent, meaning the star’s own rotation blurs any surface line across nearly a seventh of its own energy. No spectrometer resolution, however good, recovers a narrow feature from that: the line was defeated by the star’s rotation, not by any instrument’s limitations, decades before anyone built the calorimeter that could have resolved it if the star had been slower. That is the wall the paper’s audit keeps running into. Rotational broadening formalism for exactly this kind of feature was worked out in full general relativity by Peter Chang, Sharon Morsink, Lars Bildsten, and Ira Wasserman, who showed even before the 552-hertz spin was measured that a star spinning near 300 to 600 hertz would statistically be expected to show a feature as deep as the one originally claimed only five to twenty percent of the time [8].

A cryostat dewar for an X-ray microcalorimeter on an integration stand, its forward aperture door swung half open and a gold thermal blanket flap hanging loose beside the opening
Figure 2. Getting an array down to the millikelvin stage a dewar like this is built for is a mechanical problem nearly as demanding as the physics it exists to resolve, and the cooling chain behind it is exactly what a descoped, current-generation X-ray spectrometer had to redesign from scratch [@barret-2018].

The obvious fix — find a slow rotator — runs into a second wall

If rotation is the problem, the obvious answer is to look at neutron stars that barely rotate at all. The arithmetic supports this completely: a star spinning at roughly one hertz smears a surface line by only a few parts in a hundred thousand, and several real, slowly rotating accreting neutron stars sit far below any spectrometer’s resolution floor purely from the rotation term. The trouble is that the slowest rotators in the observable population are almost all magnetically confined accretors, and the same magnetic field that has spun them down over their lifetime also displaces any line they show by an amount that swamps a gravitational-redshift signal outright. Her X-1, spinning at a leisurely 1.24 seconds, clears the rotational wall by three orders of magnitude — and shows a proton-cyclotron line at 37.4 keV that implies a surface field near 3×10123\times10^{12} gauss, a line energy that itself shifts by more than six percent whenever the source’s accretion flux doubles. 4U 1626-67, spinning even more slowly at 7.66 seconds, carries a comparable multi-teragauss field and its own roughly 36-keV cyclotron feature. Slow rotation and a clean, field-free surface turn out to be almost mutually exclusive in the known population: the property that lets a star escape the rotational wall is, in essentially every observed case, produced by the same magnetic field that then contaminates the very redshift the line was supposed to measure.

The narrow exception the paper’s authors point to is a specific, unglamorous class of object: central compact objects, the so-called anti-magnetar remnants left behind in some supernova remnants, whose dipole fields are inferred to be unusually weak despite showing reported surface absorption features. The archetype, 1E 1207.4-5209, is a slow rotator by the same rotational-smearing arithmetic as Her X-1 and 4U 1626-67, but without the strong-field problem that disqualifies them — the one class in the current shortlist that plausibly clears both the rotational wall and the magnetic-contamination floor at once, though its detailed parameters are adopted from the literature rather than independently re-derived in the paper. It is a short list for a specific reason: slow rotation and a weak field are almost never found in the same star, and the sources where they are both true are correspondingly rare.

Today’s bottleneck is the standard, not the spectrometer

Here the paper reaches its least comfortable conclusion. Modern X-ray calorimeters are already far more precise than the equivalence-principle test needs. XRISM’s Resolve instrument delivers an energy resolution near 4.5 eV at 6 keV in orbit — a fractional precision of about 7.5×10−47.5\times10^{-4} — while the still-descoped NewAthena X-IFU is specified at 4 eV up to 7 keV, a comparable 5.7×10−45.7\times10^{-4}, itself a relaxation from the instrument’s original 2.5 eV design target before a 2022-23 cost-driven mission reformulation [14]. Both figures comfortably clear what a line comparison needs. What does not clear the bar is the standard the line would be compared against: at today’s roughly ten-percent NICER radius precision, the achievable bound on the equivalence-principle couplings floors out regardless of how sharp the line is, and a line resolved to better than about two percent is, by the paper’s own arithmetic, wasted precision until the redshift standard itself improves. Driving that radius uncertainty down to three percent, and then to one percent, is what actually moves the bound — not a better spectrometer.

ADVERTISEMENT

That inverts the obvious instinct that a line-search program is primarily an instrument problem. It is, instead, substantially a mass-radius problem: the fastest realistic route to a stronger bound is not building a sharper calorimeter, since XRISM and NewAthena-class instruments already outperform what today’s standard can use, but tightening the NICER radius measurement of whichever source eventually shows a credible line — continuing exactly the observational program the NICER collaboration is already running for equation-of-state reasons on sources like PSR J0437-4715 [3]. A calorimeter proposal that does not also secure a better mass-radius measurement of its target is, on this accounting, proposing half of the necessary experiment.

A microcalorimeter detector array in its gold-plated mounting flange, held above a clean-bench under a laminar-flow hood with one alignment pin not yet engaged
Figure 1. A calorimeter-class array like this one already resolves a line to a few parts in ten thousand — better than XRISM's Resolve instrument needs for the bound this paper describes — yet its own in-orbit resolution near 6 keV still has to be protected against high-count-rate degradation before that precision is real [@mizumoto-2025-ii].

2026’s nearest attempt does the inversion backward, and that is the point

The clearest illustration of why the ordering matters arrived while the paper was being finished. Roberto Iaria and collaborators reported, in data taken hours after a carbon superburst on the ultracompact binary 4U 1820-30, an absorption feature at 3.8±0.053.8\pm0.05 keV, with a width and equivalent width consistent with a real spectral line rather than noise, at a significance the authors themselves describe across their observations as ranging from 2.5 to 8 sigma [10]. Interpreted as a gravitationally redshifted iron line, the feature implies 1+z=1.72±0.051+z=1.72\pm0.05↗, a compactness near 0.33, and — worked through the same exterior relation as any other neutron star — a mass of roughly 1.8 to 2.3 solar masses and a radius of roughly 8.3 to 11 kilometres. The authors describe their own interpretation as still tentative, and they make no equivalence-principle claim about it at all.

That last point is the one worth sitting with. Because 4U 1820-30 has no independent NICER pulse-profile mass-radius measurement, Iaria and collaborators had no external redshift standard to check their line against — so the single feature had to supply both the redshift and the compactness at once, exactly the circular construction that undid the EXO 0748-676 claim two decades earlier. It is not a counterexample to the paper’s argument; it is close to a live demonstration of it, showing that even with a modern instrument and a highly capable team, the historical order of operations — let the line measure the star — persists whenever an independent standard does not exist for the source in question. A NICER-class mass-radius measurement of 4U 1820-30, run before or alongside any deeper follow-up of this feature, would convert the same line from a compactness estimate into the kind of physics test the rest of this paper is built to make possible.

The measurements underlying that redshift standard belong to the NICER collaboration — Miller, Lamb, and colleagues; Riley, Watts, and colleagues; Choudhury, Salmi, and colleagues — and the equation-of-state machinery built on top of them belongs to the wider nuclear-astrophysics community. The EXO 0748-676 story, from Cottam, Paerels, and Mendez’s original claim through Galloway and collaborators’ spin measurement to Lin, Özel, Chakrabarty, and Psaltis’s incompatibility argument, is the field’s own hard-won verdict, not a retrospective judgment imposed by this paper. A structurally similar redshift-as-test logic already exists for white dwarfs — Sirius B’s gravitational redshift, measured differentially against its companion star to 80.65±0.7780.65\pm0.77 km/s and cross-checked against an independent dynamical mass to 2.5 percent [12] — though at a potential more than three orders of magnitude shallower than any neutron-star surface. What the new paper adds is the packaging: the standard-candle inversion itself, the bound algebra connecting a line discrepancy to a declared model of equivalence-principle-violating couplings, the audit of every claimed surface feature against that framework, the corrected rotational-broadening coefficient, and the observing case showing that the redshift standard, not the spectrometer, is what currently limits the whole program.

Read the full paper (PDF)