How Constitutional AI Actually Constrains a Model's Behavior
Constitutional AI keeps RLHF's reward-model-and-reinforcement-learning core intact and swaps the source of one training dataset, a document standing in for a person's judgment. That swap is narrower, and its limits more specific, than the name suggests.