
A new model of alignment shows why even near-perfect training can still ship a catastrophic AI
One of the longest-running worries in AI safety circles holds that human values are fragile: optimize too hard for an imperfect stand-in of what people actually want, and the result is a world nobody wanted. A new preprint seeks to put that intuition on formal footing, and its conclusions are sobering. Even when a hypothetical […]










