A statistical test built for repeated looks finds DESI’s dark energy signal fragile

The strongest hint so far that dark energy changes with time may rest on a single slice of galaxy data, according to a reanalysis built around a statistical method designed for surveys that get re-tested release after release.

The Dark Energy Spectroscopic Instrument (DESI) released its second batch of baryon acoustic oscillation measurements, and the combination of those data with supernova and cosmic microwave background results pointed to a time-varying dark energy equation of state at a claimed significance of roughly 3 to 4 sigma, depending on which supernova catalog was included. Taken at face value, that would be the first statistically significant evidence that dark energy is not a cosmological constant.

The new analysis, by Jinyoung Kim of Stanford University and Two Sigma Investments, David F. Mota of the University of Oslo, and Andrius Tamosiunas of Oslo and Case Western Reserve University, applies an e-process to the same DR1 to DR2 BAO sequence. An e-process is a likelihood-ratio based measure of evidence whose validity does not depend on how often the question is revisited. An e-value has an expectation of at most one under the null hypothesis, so a large observed value counts as evidence against it; rejecting when the value reaches 1 divided by the chosen false-detection rate keeps that rate bounded. Each DESI release invites a fresh test of whether dark energy is consistent with the cosmological constant, and standard Wilks-based sigma values are interpreted as if the test had been run only once. Repeating the same test at DR1, DR2, DR3 and beyond inflates the chance that a fluctuation looks significant. The e-process controls the false-detection probability under continued testing, with no trials-factor penalty for the number of releases. The authors describe it as the first application of anytime-valid sequential inference to a cosmological dataset.

Under a pre-specified alternative concentrated near the direction the DR2 data prefer, the running evidence reaches 33.97 at DR2, crossing the threshold of 20 that corresponds to an illustrative 5% false-detection rate. But the evidence is not spread across the survey. A leave-one-out decomposition places most of it, 78.6%, in a single redshift bin, LRG2, at an effective redshift of 0.706. Remove that bin and the running e-value collapses to 0.49, which mildly favors the cosmological constant. Allow any of the seven DESI bins to have produced the excess, and the signal does not survive a look-elsewhere correction: the per-bin product is 4.46 and the arithmetic mean 1.52, both well below the threshold.

If our reporting has earned your trust, consider helping us continue our work.

Keep quality journalism alive

The verdict also depends on what the test was designed to detect. Narrow priors placed near the data’s preferred region reject the cosmological constant, while wide ones do not. A physically motivated thawing-quintessence alternative rejects more strongly than any agnostic choice, with a running value near 1,100, while freezing-quintessence models fall below the threshold. Adding the compressed Planck cosmic microwave background likelihood to the BAO data drops the default e-value to 2.19, below threshold.

The authors argue the concentration is difficult to reconcile with a smoothly evolving equation of state. In simulations of a genuinely smooth signal, LRG2 alone would carry the observed share of the evidence in only 3.2% of realisations, and the authors say the pattern is consistent with an outlier or a bin-specific systematic, though their bootstrap models only the published Gaussian covariance and cannot exclude such a systematic directly.

Several caveats qualify the result. The background cosmology is held fixed at Planck 2018 values, and the headline e-values move by two orders of magnitude across self-consistent Planck columns, with one column falling below the rejection threshold. Every number conditions on the Chevallier-Polarski-Linder parametrization of the dark energy equation of state, which the authors flag as the most direct robustness test still open. The authors also recommend that future releases from DESI, Euclid, Roman and LSST report an anytime-valid e-process value, with its test specification stated, alongside conventional sigma significances.

What could settle the question: simulations show that halving the LRG2 error would push the e-value well past threshold if evolving dark energy is real, and leave it near zero if the cosmological constant is correct. A decisive third DESI release therefore needs sharper measurements of the bins the survey already has, above all LRG2, rather than only new high-redshift ones. Whether that bin reflects new physics or a localised systematic, the authors write, will be resolved only by additional data.

Sources

Scroll to Top