› working / artifacts
capstone
Introspection Fine-Tuning
Replicated IFT, developed an improved objective function, and fitted jacobian lenses on Llama 3.2 and Qwen3.
project
Harmlessness Finetuning
Finetuning Llama 3.2 1B with SFT and DPO + LoRA for improved harmlessness.
bootcamp
Cambridge Alignment Bootcamp
My work through the ARENA curriculum in interpretability and alignment as part of CBAI's CAMBRIA program.
writing
Corporate Orthogonality
Structural misalignment and the limit of technical approaches.
research
Dimensionality Reduction
A comparison of PCA and linear autoencoders on MNIST.
statement
What are we building for?
Artificial intelligence, alignment, and building a future worth wanting.
essay
Input-Attribution Methods Cannot Justify AI Decisions
Mechanistic interpretability as the path forward towards achieving a duty of justification.
essay
We Have Yet to Cut Off the Head of the King
A misdiagnosis of power and Foucault's challenge to domination theory.