What is Alignment?

Imagine an insurance company builds an assistant to help process claims. The customer wants a fast answer. The claims team wants enough evidence to make a sound decision. Legal wants a process the company can defend. The public wants people treated consistently.

Which one of those goals should the assistant follow? Ideally, some negotiated combination of all of them. AI alignment is the work of making an AI system pursue the intended goals, follow relevant constraints, and remain under meaningful human direction.

Whose Goal?

Set the target for an insurance claim assistant

Pressure for speed70%
Pressure for caution67%
Unresolved tension4%

Weighted target = stakeholder influence × preferred behavior. The arithmetic is easy. Justifying the weights is the alignment work.

Alignment starts with a target that belongs to real stakeholders. Changing who has influence changes the behavior we call aligned.

Alignment has at least three layers. We need to decide what behavior we want, translate that behavior into training signals and system requirements, and verify that the behavior survives deployment. A failure at any layer can produce a capable system that does the wrong thing.

This makes alignment a sociotechnical problem. The model matters. So do the interface, data, reward, tools, review process, business incentives, and people affected by the decision.

The TV show "Silicon Valley" discusses what happens when alignment goes wrong (spoiler alert)

A Working Definition

An AI system is aligned when its behavior reliably advances the intended human purpose, respects the constraints that matter in context, and supports correction or intervention when conditions change.

Checkpoint

Why can alignment rarely be reduced to maximizing one stakeholder's preferred metric?