
The alignment problem in plain terms
An AGI system optimizing for a goal it was given imperfectly can cause serious harm — not from malice, but from precision. That is the alignment problem. This program treats it as an engineering challenge, not a science fiction scenario.
We cover reward modeling, constitutional AI, debate-based alignment, and interpretability tooling. You will run experiments with smaller-scale models to observe misalignment failure modes directly, which makes the theory considerably easier to retain.
Practical lab component
Four of the eight sessions include hands-on labs using open-source frameworks. You will instrument a reward model, observe reward hacking in a controlled environment, and test an interpretability probe on a transformer layer. These are not toy exercises — they reflect methods used in current alignment research.
Realistic expectations
Alignment research is genuinely hard and unsolved. This program does not promise you will leave with answers. It gives you the vocabulary, tooling, and mental models to contribute meaningfully to the conversation — whether in a research role, a policy context, or as an engineer building safety-critical AI systems.
AGI Safety and Alignment: Technical Approaches
The alignment problem in plain terms
An AGI system optimizing for a goal it was given imperfectly can cause serious harm — not from malice, but from precision. That is the alignment problem. This program treats it as an engineering challenge, not a science fiction scenario.
We cover reward modeling, constitutional AI, debate-based alignment, and interpretability tooling. You will run experiments with smaller-scale models to observe misalignment failure modes directly, which makes the theory considerably easier to retain.
Practical lab component
Four of the eight sessions include hands-on labs using open-source frameworks. You will instrument a reward model, observe reward hacking in a controlled environment, and test an interpretability probe on a transformer layer. These are not toy exercises — they reflect methods used in current alignment research.
Realistic expectations
Alignment research is genuinely hard and unsolved. This program does not promise you will leave with answers. It gives you the vocabulary, tooling, and mental models to contribute meaningfully to the conversation — whether in a research role, a policy context, or as an engineer building safety-critical AI systems.
Program structure
What gets covered and in what order — no filler, no repetition.
Program Outline
- Module 1
- Failure modes of reward specification — Goodhart's Law in practice
- Module 2
- Reinforcement learning from human feedback (RLHF): mechanics and limitations
- Module 3
- Constitutional AI and rule-based constraint systems
- Module 4
- Lab: reward hacking demonstration with a gridworld agent
- Module 5
- Interpretability methods: probing, activation patching, causal tracing
- Module 6
- Lab: building a simple linear probe on a language model
- Module 7
- Scalable oversight and debate as alignment strategies
- Module 8
- Capstone: written alignment proposal for a specified AGI use case
Ready to start?
Seats fill up quickly — once the cohort closes, the next opening is months away.