
One of the longest-running worries in AI safety circles holds that human values are fragile: optimize too hard for an imperfect stand-in of what people actually want, and the result is a world nobody wanted. A new preprint seeks to put that intuition on formal footing, and its conclusions are sobering. Even when a hypothetical training process guarantees that an agent’s value function closely matches human values by several different measures, there remain conditions under which a catastrophically misaligned agent passes the test and gets deployed.
The paper, Fragility of Value under Imperfect Alignment by Winter Cross of Dovetail Research, appeared on arXiv on July 30 and runs 24 pages. It models a group of researchers training a powerful agent, with a training process that guarantees the agent’s value function meets a proxy condition, a criterion meant to ensure similarity to human values. If the agent’s value function satisfies the condition, the researchers consider it safe and deploy it to optimize the world. The central question is whether catastrophic value functions can slip through. The authors define a value function as eta-catastrophic when optimization is guaranteed to drive the expected value of human welfare below a threshold eta as optimizing power grows without bound.
The paper works through three frameworks, each with a different world model and alignment technique. In the finite framework, the training process bounds the disagreement rate between the agent’s and humans’ value functions at beta. The main result is that if no state is valued as perfect by the humans, a catastrophic proxy exists whenever beta reaches one over the number of states. In practical terms, any misspecification at all admits a catastrophic value function, because the only proxy condition strict enough to exclude them is one so tight that only the true value function passes. If humans do value some states as perfect, the threshold loosens, but only in proportion to how many states they hold perfect.
The continuous framework models the world as a bounded continuous space and bounds the probability that the agent’s and humans’ values disagree by more than a tolerance alpha. Here the result is bleaker: for every possible human value function, and every pair of positive tolerances alpha and beta, a catastrophic proxy always exists. The proof constructs one by adding a narrow peak of high value at the state humans value least, narrow enough to pass the proxy condition but sufficient to pull all optimization toward the worst outcome. No matter how strict the test, the authors show, some catastrophic function passes it.
The third framework models human values as a finite list of attributes, and the training guarantee is strong: the agent responds to every attribute the humans respond to. Even then, the paper finds, a proxy can be catastrophic if it trades off between those attributes differently than humans do. The condition turns on the geometry of the feasible region of world states: if the humans value some Pareto-optimal point at or below eta, then some tradeoff proxy that cares about every attribute will optimize straight toward it. The authors extend an earlier line of research by Zhuang and Hadfield-Menell, which showed that proxies missing an attribute can be catastrophic, to demonstrate that proxies containing all attributes can be too.
The authors are careful about what the results do and do not establish. They show that catastrophic proxies exist under broad conditions, not that they are likely. Estimating probability would require a prior over proxy value functions, which the model does not supply. Still, they argue, the persistence of catastrophic outcomes even in toy models with strong alignment guarantees suggests the challenge is structural rather than a fixable training deficiency. The paper’s constructive recommendation is to favor AI designs that hold optimization pressure in check, such as quantilizers that select actions from a bounded set of top options, instead of depending on pre-deployment training alone to guarantee safety.
Sources: Fragility of Value under Imperfect Alignment (arXiv:2607.28881) (arXiv, Jul 30, 2026); Fragility of Value under Imperfect Alignment, full text (arXiv, Jul 30, 2026)

