Constitutional AI
Constitutional AI provides explicit principles the model can use to critique and revise responses, then generate preference data for further training.
The constitution moves human input upstream. People select the principles, examples, and conflict rules. The model then applies them at scale through self-critique and AI feedback.
This improves transparency compared with an unnamed preference target. It also concentrates attention on representation: whose principles appear, how specific they are, and what happens when privacy, helpfulness, and autonomy point in different directions.
💭Reflection
What principle would you add to the health assistant's constitution, and what behavior should that principle change?

