Constitutional AI
FreeSafer model alignment.
About Constitutional AI
Constitutional AI is a method for training harmless AI assistants through self-improvement without requiring human labels identifying harmful outputs. The only human oversight is a list of rules or principles, referred to as a 'constitution'. The process involves two phases: a supervised learning (SL) phase where an initial model generates self-critiques and revisions, then is finetuned on the revised responses; and a reinforcement learning (RL) phase that uses AI feedback to train a preference model, enabling RL from AI Feedback (RLAIF). The resulting assistant is harmless yet non-evasive, engaging with harmful queries by explaining its objections. Both phases can leverage chain-of-thought reasoning to improve transparency and human-judged performance. The method enables more precise control of AI behavior with far fewer human labels.
Key Features
Pros & Cons
- Dramatically reduces reliance on human labeling for harmlessness
- Produces a non-evasive AI that engages with harmful queries respectfully
- Chain-of-thought reasoning enhances decision transparency
- Allows precise control over AI behavior using a small set of principles
- Combines supervised and reinforcement learning for robust alignment
- Requires careful design of the constitutional principles
- Dependent on quality of the initial model and AI feedback models
- Two-phase training may be computationally expensive
- Effectiveness may vary across different types of harmful content