← working
bootcamp

Cambridge Alignment Bootcamp

My work through chapters 1 (transformer interpretability) and 4 (alignment science) of the ARENA curriculum as part of the Cambridge Bootcamp for Research in Interpretability and Alignment, an ML upskilling bootcamp for AI safety run by the Cambridge Boston Alignment Initiative.

Topics include building a transformer from scratch, working with TransformerLens, linear probing, sparse autoencoders, emergent misalignment, the science of misalignment, LLM psychology, and persona vectors.