Mechanistic interpretability researcher at Anthropic, working on superposition, dictionary learning and circuit tracing to understand what actually happens inside large models.
Mechanistic interpretability researcher at Anthropic, working on superposition, dictionary learning and circuit tracing to understand what actually happens inside large models.