
- Personas & Model Organisms
- Interpretability
Bio
I am the Executive Director of Principia and am concurrently in the final year of my PhD at University College London, supervised by Andrew Saxe. At Principia, we develop theoretical model organisms that offer clear insights into why neural networks learn and behave the way they do. I was previously a Pivotal Fellow, mentored by Jesse Hoogland.
Projects
Research Direction: learning dynamics, theoretical model organism
Current project ideas include, but are not limited to:
- How different post-training algorithms (e.g., SFT and RL) interact with what is learned during pretraining
- A principled understanding of when interventions during training are most effective
We will discuss a few promising directions and jointly decide what to work on. Our approach is often to design a simplified toy model or task whose inner workings we can clearly understand. Ideally, these insights then allow us to make predictions that can be tested on larger systems. To get a better sense of Principia’s research flavour and the frameworks we use, please see:
[1] https://arxiv.org/abs/1810.10531
[2] https://proceedings.mlr.press/v162/saxe22a/saxe22a.pdf
What I'm looking for in a Mentee
We are looking for a fellow with a strong analytical mindset, often with a background in physics, computational neuroscience, or an adjacent field. Experience with and understanding of LLMs, especially post-training, is a big plus.
What I'm Like as a Mentor
We will aim for a one-hour weekly meeting, adjusting the frequency as needed. I expect mentees to participate in Principia’s research meetings and events and to actively engage and collaborate with our research scientists. The fellowship project will also likely involve collaboration with other Principia researchers.