Portrait of Eyon Jang
  • Science of Model Behavior
  • Alignment Auditing
  • Scalable Oversight
  • RSI Safety
  • Multi-agent Safety
  • Misalignment generalization

Co-Mentor: Shi Feng

Eyon Jang

Research Scientist, Preparedness team, Scale AI

Bio

My research focuses on building a rigorous science of model behavior, developing reliable evaluations of frontier risk, and studying safety post-training and alignment techniques that remain effective as AI systems become more capable.

Previously, I was a MATS scholar researching exploration hacking and automated alignment auditing, mentored by David Lindner, Roland Zimmermann (Google DeepMind AGI Safety and Alignment), and Scott Emmons (Anthropic Alignment Science).

Before moving into technical AI safety, I spent 6 years as a quantitative researcher on Wall Street. I received my MSc in Statistics and Machine Learning (with Distinction) from the University of Oxford, where I was supervised by Prof. Yee Whye Teh and Prof. Benjamin Bloem-Reddy.

Projects

Research Direction: Alignment auditing and its automation

We will generally work on alignment auditing. By auditing, we broadly refer to answering fuzzy questions about a model's behavior: what are its motivations behind a misaligned behavior, and what interventions would present such behavior? Practically, being able to answer these questions and monitor latent processes such as intent is directly useful for reducing real-world harm. Academically, we are in dire need for an empirical science for alignment as we approach RSI.

We are interested in questions on two levels.

  • On the meta level, we are interested in the science of alignment auditing. What's the right target? How should we measure progress? What are threat models in automating it? See interpretive debate for a related discussion.
  • On the object level, we will build alignment auditing environments, train model organisms, and study auditing tools including white-box, training-based methods for belief editing, e.g., grafting.

There is space for many different types of work:

  • Conceptual work on designing auditing games.
  • Engineering-heavy work such as building cheating propensity evals.
  • Developing training recipes for model organism and auditing tools.

We are very focused on this direction, but welcome project proposals that are inline with our theory of change.

What we’re looking for in a Mentee

  • Familiar with broader alignment research landscape and theory of change.
  • High reasoning transparency. Cares about science and rigor in empirical work.
  • Strong conceptual and critical thinking skills.
  • Comfortable with frequent switches between fuzzy conceptual thinking and engineering work.

What we’re like as Mentors

'- I'm happy to be hands-on when that's helpful.

  • For fellowships, I usually pitch a few (~3) projects under a common research agenda and let mentees to choose between them or propose projects that are compatible with the broader theory of change.
  • I usually prefer multiple mentees to work together on projects over projects by individual mentees.
  • I generally try to make myself available for quick ~15min ad-hoc chats. I usually respond to Slack messages quickly. I enjoy structured communication with clear expectations to reduce the overhead of context loading.