
Bio
I'm a member of technical staff at Apollo Research, where I try to build a science of scheming, most recently studying how long-horizon RL can degrade corrigibility. Before joining Apollo, I participated in MATS and LASR, where I worked on scheming propensity evaluations, honeypots, and automated auditing.
Projects
I'm excited about supervising projects related to Apollo's science of scheming agenda, where we study scheming and the training dynamics that give rise to scheming-related propensities. In particular, I'm interested in what kinds of training make models develop goals, and when those goals lead them to subvert oversight.
Most threat models for catastrophic outcomes assume at some point that the model "has a goal", but we don't understand well how training gives models goals at all, or when those goals take the more concerning forms: ambitious, broadly scoped, beyond-episode, real-world facing.
We aren't yet in a training regime that strongly and directly incentivizes instrumentally convergent goals, but this should change as labs train on longer-horizon tasks or adopt continual learning, which they are heavily incentivized to do and explicitly plan to.
To study this, we would train model organisms on proxy tasks for these future regimes and identify the conditions under which models crystallize goals, and whether these goals motivate them to subvert human oversight when they conflict with their principals' goals. Day to day, this means forming hypotheses, training model organisms, and building evaluations to test them.
A strong result by the end of April would be a model organism that comes to care strongly about a goal through completely benign training, on a proxy task we think isolates a concerning aspect of future training regimes. Ablations would show which training conditions cause this, and we'd write it up as a paper or LessWrong post.
The exact project isn't fixed though. I want my fellow and me to work on something we're both excited about and think is most important when the program starts.