AI Safety
The Main Technical Approaches to AI Alignment, Compared
RLHF, Constitutional AI, debate, iterated amplification, weak-to-strong generalization, and verifiable-reward RL are six distinct, documented bets on how oversight can survive a system that starts to outgrow the people evaluating it.