A deeply researched exploration of how we teach machines right from wrong — and how that effort mirrors our own moral development. Christian traces the alignment problem from early RLHF experiments to fairness audits, framing the entire field through the lens of what it means to align a system with human values. Particularly relevant to my work in interpretability: the question of whether a model has truly learned something versus whether it has learned to appear as though it has is exactly the gap my unlearning research probes.