Alignment & Safety
What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Reinforcement learning from human feedback does not install correctness. It fits a scalar model of what raters approved of and then pushes a policy up that surface — which is a different thing, in identifiable places.