A (Very) Short Introduction to the Alignment Problem
AI alignment steers artificial intelligence systems towards behaving according to human values and preferences. This article explores how limits in generalization, value specification, interpretability, scalable oversight, and verification collectively make perfect alignment difficult to achieve.