Bio
I’m an interpretability researcher, employed by Adecco, supporting Google DeepMind in the Language Model Interpretability team. Previously, I worked at FAR.AI, Apollo Research, completed a PhD in Astronomy and Neel Nanda’s MATS program. My research interests lie mostly in fundamental mech interp, new methods & good toy models. You can find some of my takes in my short-form comments. I try to focus on neglected ideas, which just means I avoid doing plain SAE projects. You can find my past projects on LessWrong as well as arXiv. I’ve mentored around 9 projects across various programs over the past year [1,2,3,4,5, more in preparation], including LASR, SPAR, MARS, and Pivotal.
Projects
Research Direction: Fundamental mechanistic interpretability (activation plateaus, computation in superposition)
I'm planning various projects in the activation plateau and computation in superposition direction.
Background: Common sparse dictionary learning techniques (e.g. SAEs) find features as the vectors that optimally compress a dataset of activations. I claim that dataset structure dominates: SAEs trained on randomly initialized transformers look interpretable, so "looking interpretable" is little evidence. Instead, I think it's a good bet to focus on "poking the model": measure how the model responds to changes in internals (activation perturbations, weight perturbations, etc.) on individual prompts. This direction seems neglected, avoids the dataset issues, and initial results (activation plateaus) seem promising. If such a technique finds interpretable concepts, that's a coincidence: the technique had no access to what we humans would have wanted to see, so these concepts must come from the model.
- Reverse-engineer activation plateaus. LLMs are highly robust to activation perturbations up to a certain magnitude, but change quickly afterwards, and only for real activations. How do ~1B MLP weights encode >>1B points in activation space? Find whatever rule the MLPs encode, then test whether it predicts plateaus.
- The shape of plateaus. Are activation plateaus spheres of robustness, or tubes / manifolds within which we can get smooth behaviour? We have tentative evidence for plateaus being shorter in feature directions than in mixed directions. Can we find a dictionary of features based on the location of plateaus only?
- Computation in superposition. LLMs seemingly perform more computations than naively possible. Test activation plateaus in toy models of Computation in Superposition and Error Correction, and test which phenomenological effects correspond to which part of the theory.
Almost all of these phenomena appear in relatively small LLMs or even toy models, making larger numbers of experiments easy.
A strong result: We find a specific rule that predicts plateaus (e.g. activations equal to the sum of some dictionary elements) that advances the field of interpretability. Written up as a paper & blog post, posted publicly (on LessWrong and arXiv).
Note: My projects are likely more open-ended than typical projects, and won't necessarily result in a workshop or conference publication. Let me know (in interviews) if this is important for you.
What I'm looking for in a Mentee
I'm especially keen for applications from experienced researchers (e.g. postdocs) from other fields, OR applicants that have already thought about interpretability and maybe done a project in their free time. I expect all applicants to be proficient in Python and Pytorch, and to have spent >40h working with transformers or reading & thinking about interpretability. I expect mentees to be proactive about next steps (e.g. think of and run follow-up experiments before our meeting), conscientious (especially: red-teaming your own results, looking for ways in which our results could be misleading), and critical (challenge my plans, push back on ideas you think are bad).
What I'm Like as a Mentor
This year I'm planning to mentor individual mentees rather than teams, with a weekly 1h call to answer questions, discuss progress, and give advice on the next steps. Between calls we communicate via Slack, and I typically can respond to questions within 24h. Please communicate frequently and early (e.g. send screenshots of your Colab notebook with a line of commentary); no mentee has ever managed to over-communicate :)
Mentored projects
Mentored research
View All Research6 Jul 2026
Compressed Computation under L⁴ Loss is likely Computation in Superposition
Researcher: Francisco Ferreira da Silva
Mentor: Stefan Heimersheim
Technical AI Safety
3 Jul 2026
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
Researcher: Arnau Marin-Llobet
Mentor: Stefan Heimersheim
Technical AI Safety
23 Jun 2026
Evidence for Feature-Specific Error Correction in LLMs
Researcher: Francisco Ferreira da Silva
Mentor: Stefan Heimersheim
Technical AI Safety
30 Sep 2025
Activation Probes Are Reliable With a Handful of Positive Examples
Researcher: Riya Tyagi
Mentor: Stefan Heimersheim
Technical AI Safety
16 Jul 2025
Benchmarking Deception Probes via Black-to-White Performance Boosts
Researcher: Avi Parrack
Mentor: Stefan Heimersheim
Technical AI Safety
