AI Safety
Outer Alignment, Inner Alignment, and Why the Difference Matters
Aligned is not one property but two — whether the specified objective is right, and whether the trained system actually pursues it. A documented case shows exactly where the two split apart, and why neither one alone is what people mean by safe.