Unit 5
AI Alignment
How do we build AI systems that pursue intended goals, remain reliable as conditions change, and stay accountable to the people affected by them? This unit moves from reward design and preference learning to assurance, values, and governance.
Chapter 1
The Alignment Problem
You will be able to distinguish alignment from capability, trace misalignment from intent to impact, and diagnose specification gaming, reward hacking, and goal misgeneralization.
Chapter 2
Learning Human Preferences
You will be able to explain reinforcement learning, compare forms of human and AI feedback, step through RLHF, and evaluate constitutional approaches to preference learning.
Chapter 3
Learning to Generalize Beyond Training
You will be able to identify distribution shift, probe goal generalization, design robust training curricula, and reason about cooperation.
Chapter 4
Alignment Assurance
You will be able to design scoped safety evaluations, operationalize human values into testable requirements, and connect to the book's dedicated treatments of red teaming, interpretability, and monitoring.
Chapter 5
Human Values and Ethics
You will be able to compare value-aggregation rules, evaluate moral-policy approaches, reason across alignment tradeoffs, and design for pluralism without sacrificing factual or safety boundaries.
Chapter 6
Governance Overview
This is a very basic interview of AI governance. You will be able to place governance controls across the AI lifecycle and translate standards into evidence.